JIT: some ideas on high-level representation of runtime operations in IR #9056

Description

@AndyAyersMS

To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

Motivations:

  • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
  • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
  • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
  • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
  • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
  • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

Possible candidates for this kind of encapsulation include

  • Runtime lookup
  • Static field access
  • Box (already semi-encapsulated)
  • Unbox
  • Cast/Isint
  • Allocation (already encapsulated)
  • Class initialization

The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

fgMorphTree BB01, stmt 3 (before)
[000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
[000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
[000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
[000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
[000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
[000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class

By default the wrapper just evaporates in morph:

Optimizing call to Type:op_Equality to simple compare via EQ
Optimizing compare of types-from-handles to instead compare handles
fgMorphTree BB01, stmt 3 (after)
[000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
[000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

But in morph and upstream it can be used to trigger new optimizations.

category:implementation
theme:ir
skill-level:expert
cost:medium

Metadata

Metadata

Assignees

No one assigned

    Labels

    JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      JIT: some ideas on high-level representation of runtime operations in IR #9056

      Description

      @AndyAyersMS

      To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

      Motivations:

      • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
      • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
      • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
      • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
      • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
      • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

      Possible candidates for this kind of encapsulation include

      • Runtime lookup
      • Static field access
      • Box (already semi-encapsulated)
      • Unbox
      • Cast/Isint
      • Allocation (already encapsulated)
      • Class initialization

      The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

      We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

      Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

      As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

      So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

      The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

      If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

      And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

      fgMorphTree BB01, stmt 3 (before)
      [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
      [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
      [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
      [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
      [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
      [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
      

      By default the wrapper just evaporates in morph:

      Optimizing call to Type:op_Equality to simple compare via EQ
      Optimizing compare of types-from-handles to instead compare handles
      fgMorphTree BB01, stmt 3 (after)
      [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
      [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

      But in morph and upstream it can be used to trigger new optimizations.

      category:implementation
      theme:ir
      skill-level:expert
      cost:medium

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

        Type

        No type

        Projects

        No projects

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          JIT: some ideas on high-level representation of runtime operations in IR #9056

          Description

          @AndyAyersMS

          To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

          Motivations:

          • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
          • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
          • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
          • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
          • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
          • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

          Possible candidates for this kind of encapsulation include

          • Runtime lookup
          • Static field access
          • Box (already semi-encapsulated)
          • Unbox
          • Cast/Isint
          • Allocation (already encapsulated)
          • Class initialization

          The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

          We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

          Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

          As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

          So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

          The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

          If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

          And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

          fgMorphTree BB01, stmt 3 (before)
          [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
          [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
          [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
          [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
          [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
          [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
          

          By default the wrapper just evaporates in morph:

          Optimizing call to Type:op_Equality to simple compare via EQ
          Optimizing compare of types-from-handles to instead compare handles
          fgMorphTree BB01, stmt 3 (after)
          [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
          [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

          But in morph and upstream it can be used to trigger new optimizations.

          category:implementation
          theme:ir
          skill-level:expert
          cost:medium

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

            Type

            No type

            Projects

            No projects

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              JIT: some ideas on high-level representation of runtime operations in IR #9056

              Description

              @AndyAyersMS

              To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

              Motivations:

              • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
              • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
              • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
              • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
              • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
              • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

              Possible candidates for this kind of encapsulation include

              • Runtime lookup
              • Static field access
              • Box (already semi-encapsulated)
              • Unbox
              • Cast/Isint
              • Allocation (already encapsulated)
              • Class initialization

              The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

              We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

              Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

              As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

              So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

              The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

              If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

              And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

              fgMorphTree BB01, stmt 3 (before)
              [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
              [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
              [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
              [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
              [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
              [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
              

              By default the wrapper just evaporates in morph:

              Optimizing call to Type:op_Equality to simple compare via EQ
              Optimizing compare of types-from-handles to instead compare handles
              fgMorphTree BB01, stmt 3 (after)
              [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
              [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

              But in morph and upstream it can be used to trigger new optimizations.

              category:implementation
              theme:ir
              skill-level:expert
              cost:medium

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

                Type

                No type

                Projects

                No projects

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  JIT: some ideas on high-level representation of runtime operations in IR #9056

                  Description

                  @AndyAyersMS

                  To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

                  Motivations:

                  • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
                  • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
                  • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
                  • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
                  • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
                  • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

                  Possible candidates for this kind of encapsulation include

                  • Runtime lookup
                  • Static field access
                  • Box (already semi-encapsulated)
                  • Unbox
                  • Cast/Isint
                  • Allocation (already encapsulated)
                  • Class initialization

                  The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

                  We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

                  Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

                  As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

                  So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

                  The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

                  If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

                  And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

                  fgMorphTree BB01, stmt 3 (before)
                  [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
                  [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
                  [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                  [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
                  [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                  [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
                  

                  By default the wrapper just evaporates in morph:

                  Optimizing call to Type:op_Equality to simple compare via EQ
                  Optimizing compare of types-from-handles to instead compare handles
                  fgMorphTree BB01, stmt 3 (after)
                  [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
                  [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

                  But in morph and upstream it can be used to trigger new optimizations.

                  category:implementation
                  theme:ir
                  skill-level:expert
                  cost:medium

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

                    Type

                    No type

                    Projects

                    No projects

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      JIT: some ideas on high-level representation of runtime operations in IR #9056

                      Description

                      @AndyAyersMS

                      To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

                      Motivations:

                      • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
                      • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
                      • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
                      • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
                      • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
                      • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

                      Possible candidates for this kind of encapsulation include

                      • Runtime lookup
                      • Static field access
                      • Box (already semi-encapsulated)
                      • Unbox
                      • Cast/Isint
                      • Allocation (already encapsulated)
                      • Class initialization

                      The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

                      We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

                      Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

                      As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

                      So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

                      The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

                      If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

                      And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

                      fgMorphTree BB01, stmt 3 (before)
                      [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
                      [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
                      [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                      [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
                      [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                      [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
                      

                      By default the wrapper just evaporates in morph:

                      Optimizing call to Type:op_Equality to simple compare via EQ
                      Optimizing compare of types-from-handles to instead compare handles
                      fgMorphTree BB01, stmt 3 (after)
                      [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
                      [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

                      But in morph and upstream it can be used to trigger new optimizations.

                      category:implementation
                      theme:ir
                      skill-level:expert
                      cost:medium

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

                        Type

                        No type

                        Projects

                        No projects

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          JIT: some ideas on high-level representation of runtime operations in IR #9056

                          Description

                          @AndyAyersMS

                          To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

                          Motivations:

                          • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
                          • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
                          • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
                          • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
                          • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
                          • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

                          Possible candidates for this kind of encapsulation include

                          • Runtime lookup
                          • Static field access
                          • Box (already semi-encapsulated)
                          • Unbox
                          • Cast/Isint
                          • Allocation (already encapsulated)
                          • Class initialization

                          The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

                          We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

                          Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

                          As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

                          So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

                          The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

                          If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

                          And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

                          fgMorphTree BB01, stmt 3 (before)
                          [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
                          [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
                          [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                          [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
                          [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                          [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
                          

                          By default the wrapper just evaporates in morph:

                          Optimizing call to Type:op_Equality to simple compare via EQ
                          Optimizing compare of types-from-handles to instead compare handles
                          fgMorphTree BB01, stmt 3 (after)
                          [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
                          [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

                          But in morph and upstream it can be used to trigger new optimizations.

                          category:implementation
                          theme:ir
                          skill-level:expert
                          cost:medium

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

                            Type

                            No type

                            Projects

                            No projects

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              JIT: some ideas on high-level representation of runtime operations in IR #9056

                              Description

                              @AndyAyersMS

                              To better support high-level optimizations, it makes sense to try and defer or encapsulate some of the more complex runtime lowerings in the JIT IR. Here are some thoughts on the matter.

                              Motivations:

                              • High-level optimizations would prefer to see logical operators (even if complex) rather than a complex tree or tree sequence
                              • Many times these operators become dead and cleaning up after them can be complex if they’ve been expanded
                              • Sometimes these operators can be abstractly simplified if they are partially dead. For instance a box used only to feed a type test can become a type lookup.
                              • Properties of these operators are not always evident from their expansions, and the expansions can vary considerably, making “reparsing” within the jit to recover information lost during expansion problematic
                              • Often of these operators have nice properties (invariant, nonfaulting) and would be good candidates for hoisting, but their complex shape makes this difficult/costly.
                              • Often the equivalence of two such operators can be stated rather simply as equivalence of some abstract inputs, making CSE/value numbering simple.

                              Possible candidates for this kind of encapsulation include

                              • Runtime lookup
                              • Static field access
                              • Box (already semi-encapsulated)
                              • Unbox
                              • Cast/Isint
                              • Allocation (already encapsulated)
                              • Class initialization

                              The downside to encapsulation is that the subsequent expansion is context dependent. The jit would have to ensure that it could retain all the necessary bits of context so it could query the runtime when it is time to actually expand the operation. This becomes complicated when these runtime operators are created during inlining, as sometimes inlining must be abandoned when the runtime operator expansions become complex. So it could be this approach becomes somewhat costly in space (given the amount of retained context per operator) or in time (since we likely must simulate enough of the expansion during inlining to see if problematic cases arise).

                              We’d also have more kinds of operations flowing around in the IR and would need to decide when to remove/expand them. This can be done organically, removing the operations just after the last point at which some optimization is able to reason about them. Initially perhaps they’d all vanish after inlining or we could repurpose the object allocation lowering to become a more general runtime lowering.

                              Instead of full encapsulation, we might consider relying initially on partial encapsulation like we do now for box: introduce a “thin” unary encapsulation wrapper over a fully expanded tree that identifies the tree as an instance of some particular runtime operation (and possibly, as in box, keeping tabs on related upstream statements) with enough information to identify the key properties. Expansion would be simple: the wrapper would disappear at a suitable downstream phase, simply replaced by its content. These thin wrappers would not need to capture all the context, but just add a small amount of additional state. Current logic for abandoning inlines in the face of complex expansion would apply, so no new logic would be needed.

                              As opportunities arise we can then gradually convert the thin wrappers to full encapsulations; most “upstream” logic should not care that much since presumably the expanded subtrees, once built, do not play any significant role in high-level optimization, so their creation could be deferred.

                              So I’m tempted to say that thin encapsulation gives us the right set of tradeoffs, and start building upon that.

                              The likely first target is the runtime lookups feeding type equality and eventually type cast operations. Then probably static field accesses feeding devirtualization opportunities.

                              If you’re curious what this would look like, here’s a prototype: master..AndyAyersMS:WrapRuntimeLookup

                              And here’s an example using the prototype. In this case the lookup tree is split off into an earlier statement, but at the point of use we can still see some information about what type the tree intends to look up. A new jit interface call (not present in the fork above) can use this to determine if the types are possibly equal or not equal, even with runtime lookups for one or both inputs.

                              fgMorphTree BB01, stmt 3 (before)
                              [000037] --C-G------- * JTRUE void [000035] ------------ | /--* CNS_INT int 0
                              [000036] --C-G------- \--* EQ int [000032] --C-G------- \--* CALL int System.Type.op_Equality
                              [000028] --C-G------- arg0 +--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                              [000026] ------------ arg0 | \--* RUNTIMELOOKUP long 0x7ffcc5769428 class
                              [000025] ------------ | \--* LCL_VAR long V03 loc1 [000031] --C-G------- arg1 \--* CALL help ref HELPER.CORINFO_HELP_TYPEHANDLE_TO_RUNTIMETYPE
                              [000029] ------------ arg0 \--* CNS_INT(h) long 0x7ffcc5769530 class
                              

                              By default the wrapper just evaporates in morph:

                              Optimizing call to Type:op_Equality to simple compare via EQ
                              Optimizing compare of types-from-handles to instead compare handles
                              fgMorphTree BB01, stmt 3 (after)
                              [000037] ----G+------ * JTRUE void [000029] -----+------ | /--* CNS_INT(h) long 0x7ffcc5769530 class
                              [000170] J----+-N---- \--* NE int [000025] -----+------ \--* LCL_VAR long V03 loc1 

                              But in morph and upstream it can be used to trigger new optimizations.

                              category:implementation
                              theme:ir
                              skill-level:expert
                              cost:medium

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                JitUntriagedCLR JIT issues needing additional triagearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issue

                                Type

                                No type

                                Projects

                                No projects

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions