Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Added PerfScore support for Arm64 - #751

Merged
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore
Dec 12, 2019
Merged

Added PerfScore support for Arm64#751
briansull merged 1 commit into
dotnet:masterfrom
briansull:new-arm64-perfscore

Conversation

@briansull

@briansullbriansull commented Dec 11, 2019

Copy link
Copy Markdown
Contributor

Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
Compiles all of System.Private.CoreLib.dll for Arm64 with updated PerfScore numbers

@briansull

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib PTAL

@briansullbriansull added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Dec 11, 2019
@BruceForstall

Copy link
Copy Markdown
Contributor

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.
In any event, I think it's reasonable to go ahead and merge this and make any needed adjustments later.

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo (unhandled).
I'm not sure why you wouldn't leave this enabled in DEBUG builds, though I can see that it might be risky. That said, I don't know how often someone would think of re-enabling this if they were adding instructions.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is active work with adding instruction in the SIMD area. As I will be out ofthe office shortly I didn't want to block them at this time.

If Tanner and others want I can enable this DEBUG assertion to enforce that they add latencies for any new instructions.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be nice, IMO, to enforce this now while we are doing iterative work; rather than needing to go back and fixup everything later.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I created PR #810 to turn this assert on

@BruceForstallBruceForstall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few typos and nits

Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emit.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitxarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Based upon arm_cortex_a55_software_optimization_guide_v2.pdf
@briansull
briansull merged commit 41715bc into dotnet:masterDec 12, 2019
@TamarChristinaArm

Copy link
Copy Markdown
Contributor

Overall structure looks good. I haven't reviewed the actual latencies, and hope that perhaps @tannergooding or @TamarChristinaArm can do so.

Sure, I'll review them today.

// Branch Instructions
//

case IF_BI_0A: // b, bl_local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious.. what are bl_local and b_tail?

case INS_udiv:
if (id->idOpSize() == EA_4BYTE)
{
result.insThroughput = PERFSCORE_THROUGHPUT_4C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this be 3? Or are you applying an additional penalty? same for the X-form one below.

// Load/Store Instructions
//

case IF_LS_1A: // ldr, ldrsw (literal, pc relative immediate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming it's intentional that none of the loads have a latency? one does seem to be set for ldp and stp but the default value is PERFSCORE_LATENCY_ILLEGAL isn't it?

}
break;

case IF_LS_2D: // ld1 (vector - multiple structures)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should probably clarify that these are only for the 1 register case.


case IF_DV_1C: // fcmp vn, #0.0
result.insThroughput = PERFSCORE_THROUGHPUT_1C;
result.insLatency = PERFSCORE_LATENCY_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is off, I think this should be 1 as well, the explicit compare with 0.0 isn't more expensive than an arbitrary scalar.

case INS_fabs:
case INS_fneg:
result.insThroughput = PERFSCORE_THROUGHPUT_2X;
result.insLatency = (id->idOpSize() == EA_8BYTE) ? PERFSCORE_LATENCY_2C : PERFSCORE_LATENCY_3C / 2;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm why the different costs for the Q one? shouldn't this just be 4?

if ((id->idInsOpt() == INS_OPTS_2S) || (id->idInsOpt() == INS_OPTS_4S))
{
// S-form
result.insThroughput = PERFSCORE_THROUGHPUT_3C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These seem off, the guide says latency 12, throughput 1/9 for the S variant and latency 22 throughput 1/19 for the D form. These seem to be quite a bit cheaper.

{
// D-form
assert(id->idInsOpt() == INS_OPTS_2D);
result.insThroughput = PERFSCORE_THROUGHPUT_10C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

think this should be 19.

if (id->idOpSize() == EA_8BYTE)
{
// D-form
result.insThroughput = PERFSCORE_THROUGHPUT_6C;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm, the scalar and vector ones have the same latency and throughput in the guide. are these intentionally lower?

@ghostghost locked as resolved and limited conversation to collaborators Dec 11, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@briansull@BruceForstall@TamarChristinaArm@CarolEidt@tannergooding