Skip to content

JIT: more unexpected dynamic PGO schema mismatches #85856

Description

@AndyAyersMS

I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
{
continue; /* NOP */
}

So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

;; compiling as root
-----------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
-----------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
Using edge profiling
EfficientEdgeCountInstrumentor: preparing for instrumentation
New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
[0] New probe for BB05 -> BB01 [source]
[1] New probe for BB03 -> BB04 [source]
5 blocks, 2 probes (0 on critical edges)
;; as inlinee
-----------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
-----------------------------------------------------------------------------------------------------------------------------------------
BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
*************** Inline @[000773] Starting PHASE Profile incorporation
Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
Reconstructing block counts from sparse edge instrumentation
... adding known edge BB03 -> BB04: weight 65208
Could not find source block for schema entry 1 (IL offset/key 00000039)
... not solving because of the mismatch
... discarding profile count data: PGO data available, but IL did not match
Computing inlinee profile scale:

We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

Metadata

Metadata

Assignees

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
    Skip to content

    JIT: more unexpected dynamic PGO schema mismatches #85856

    Description

    @AndyAyersMS

    I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

    But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

    jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
    if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
    {
    continue; /* NOP */
    }

    So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

    ;; compiling as root
    -----------------------------------------------------------------------------------------------------------------------------------------
    BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
    -----------------------------------------------------------------------------------------------------------------------------------------
    BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
    Using edge profiling
    EfficientEdgeCountInstrumentor: preparing for instrumentation
    New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
    [0] New probe for BB05 -> BB01 [source]
    [1] New probe for BB03 -> BB04 [source]
    5 blocks, 2 probes (0 on critical edges)
    ;; as inlinee
    -----------------------------------------------------------------------------------------------------------------------------------------
    BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
    -----------------------------------------------------------------------------------------------------------------------------------------
    BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
    *************** Inline @[000773] Starting PHASE Profile incorporation
    Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
    Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
    Reconstructing block counts from sparse edge instrumentation
    ... adding known edge BB03 -> BB04: weight 65208
    Could not find source block for schema entry 1 (IL offset/key 00000039)
    ... not solving because of the mismatch
    ... discarding profile count data: PGO data available, but IL did not match
    Computing inlinee profile scale:
    

    We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

    Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

    Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

    Metadata

    Metadata

    Assignees

    Labels

    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
      Skip to content

      JIT: more unexpected dynamic PGO schema mismatches #85856

      Description

      @AndyAyersMS

      I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

      But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

      jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
      if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
      {
      continue; /* NOP */
      }

      So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

      ;; compiling as root
      -----------------------------------------------------------------------------------------------------------------------------------------
      BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
      -----------------------------------------------------------------------------------------------------------------------------------------
      BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
      Using edge profiling
      EfficientEdgeCountInstrumentor: preparing for instrumentation
      New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
      [0] New probe for BB05 -> BB01 [source]
      [1] New probe for BB03 -> BB04 [source]
      5 blocks, 2 probes (0 on critical edges)
      ;; as inlinee
      -----------------------------------------------------------------------------------------------------------------------------------------
      BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
      -----------------------------------------------------------------------------------------------------------------------------------------
      BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
      *************** Inline @[000773] Starting PHASE Profile incorporation
      Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
      Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
      Reconstructing block counts from sparse edge instrumentation
      ... adding known edge BB03 -> BB04: weight 65208
      Could not find source block for schema entry 1 (IL offset/key 00000039)
      ... not solving because of the mismatch
      ... discarding profile count data: PGO data available, but IL did not match
      Computing inlinee profile scale:
      

      We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

      Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

      Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

      Metadata

      Metadata

      Assignees

      Labels

      area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

      Type

      No type

      Projects

      No projects

        Milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
        Skip to content

        JIT: more unexpected dynamic PGO schema mismatches #85856

        Description

        @AndyAyersMS

        I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

        But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

        jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
        if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
        {
        continue; /* NOP */
        }

        So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

        ;; compiling as root
        -----------------------------------------------------------------------------------------------------------------------------------------
        BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
        -----------------------------------------------------------------------------------------------------------------------------------------
        BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
        Using edge profiling
        EfficientEdgeCountInstrumentor: preparing for instrumentation
        New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
        [0] New probe for BB05 -> BB01 [source]
        [1] New probe for BB03 -> BB04 [source]
        5 blocks, 2 probes (0 on critical edges)
        ;; as inlinee
        -----------------------------------------------------------------------------------------------------------------------------------------
        BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
        -----------------------------------------------------------------------------------------------------------------------------------------
        BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
        *************** Inline @[000773] Starting PHASE Profile incorporation
        Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
        Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
        Reconstructing block counts from sparse edge instrumentation
        ... adding known edge BB03 -> BB04: weight 65208
        Could not find source block for schema entry 1 (IL offset/key 00000039)
        ... not solving because of the mismatch
        ... discarding profile count data: PGO data available, but IL did not match
        Computing inlinee profile scale:
        

        We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

        Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

        Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

        Metadata

        Metadata

        Assignees

        Labels

        area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

        Type

        No type

        Projects

        No projects

          Milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
          Skip to content

          JIT: more unexpected dynamic PGO schema mismatches #85856

          Description

          @AndyAyersMS

          I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

          But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

          jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
          if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
          {
          continue; /* NOP */
          }

          So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

          ;; compiling as root
          -----------------------------------------------------------------------------------------------------------------------------------------
          BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
          -----------------------------------------------------------------------------------------------------------------------------------------
          BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
          Using edge profiling
          EfficientEdgeCountInstrumentor: preparing for instrumentation
          New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
          [0] New probe for BB05 -> BB01 [source]
          [1] New probe for BB03 -> BB04 [source]
          5 blocks, 2 probes (0 on critical edges)
          ;; as inlinee
          -----------------------------------------------------------------------------------------------------------------------------------------
          BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
          -----------------------------------------------------------------------------------------------------------------------------------------
          BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
          *************** Inline @[000773] Starting PHASE Profile incorporation
          Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
          Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
          Reconstructing block counts from sparse edge instrumentation
          ... adding known edge BB03 -> BB04: weight 65208
          Could not find source block for schema entry 1 (IL offset/key 00000039)
          ... not solving because of the mismatch
          ... discarding profile count data: PGO data available, but IL did not match
          Computing inlinee profile scale:
          

          We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

          Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

          Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

          Metadata

          Metadata

          Assignees

          Labels

          area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

          Type

          No type

          Projects

          No projects

            Milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
            Skip to content

            JIT: more unexpected dynamic PGO schema mismatches #85856

            Description

            @AndyAyersMS

            I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

            But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

            jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
            if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
            {
            continue; /* NOP */
            }

            So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

            ;; compiling as root
            -----------------------------------------------------------------------------------------------------------------------------------------
            BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
            -----------------------------------------------------------------------------------------------------------------------------------------
            BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
            Using edge profiling
            EfficientEdgeCountInstrumentor: preparing for instrumentation
            New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
            [0] New probe for BB05 -> BB01 [source]
            [1] New probe for BB03 -> BB04 [source]
            5 blocks, 2 probes (0 on critical edges)
            ;; as inlinee
            -----------------------------------------------------------------------------------------------------------------------------------------
            BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
            -----------------------------------------------------------------------------------------------------------------------------------------
            BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
            *************** Inline @[000773] Starting PHASE Profile incorporation
            Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
            Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
            Reconstructing block counts from sparse edge instrumentation
            ... adding known edge BB03 -> BB04: weight 65208
            Could not find source block for schema entry 1 (IL offset/key 00000039)
            ... not solving because of the mismatch
            ... discarding profile count data: PGO data available, but IL did not match
            Computing inlinee profile scale:
            

            We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

            Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

            Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

            Metadata

            Metadata

            Assignees

            Labels

            area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

            Type

            No type

            Projects

            No projects

              Milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
              Skip to content

              JIT: more unexpected dynamic PGO schema mismatches #85856

              Description

              @AndyAyersMS

              I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

              But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

              jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
              if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
              {
              continue; /* NOP */
              }

              So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

              ;; compiling as root
              -----------------------------------------------------------------------------------------------------------------------------------------
              BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
              -----------------------------------------------------------------------------------------------------------------------------------------
              BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
              Using edge profiling
              EfficientEdgeCountInstrumentor: preparing for instrumentation
              New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
              [0] New probe for BB05 -> BB01 [source]
              [1] New probe for BB03 -> BB04 [source]
              5 blocks, 2 probes (0 on critical edges)
              ;; as inlinee
              -----------------------------------------------------------------------------------------------------------------------------------------
              BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
              -----------------------------------------------------------------------------------------------------------------------------------------
              BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
              *************** Inline @[000773] Starting PHASE Profile incorporation
              Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
              Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
              Reconstructing block counts from sparse edge instrumentation
              ... adding known edge BB03 -> BB04: weight 65208
              Could not find source block for schema entry 1 (IL offset/key 00000039)
              ... not solving because of the mismatch
              ... discarding profile count data: PGO data available, but IL did not match
              Computing inlinee profile scale:
              

              We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

              Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

              Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

              Metadata

              Metadata

              Assignees

              Labels

              area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

              Type

              No type

              Projects

              No projects

                Milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); JIT: more unexpected dynamic PGO schema mismatches · Issue #85856 · dotnet/runtime · GitHub
                Skip to content

                JIT: more unexpected dynamic PGO schema mismatches #85856

                Description

                @AndyAyersMS

                I had hoped that #85805 would have fixed all the cases where we have dynamic PGO data available but can't match up the schema with the current flow graph and so throw away perfectly good PGO data.

                But that's not the case. The initial flow graph for a method may be slightly different, depending on whether or not the method is an inlinee. One culprit is this bit of code (there may be more I haven't spotted yet):

                jmpDist = (sz == 1) ? getI1LittleEndian(codeAddr) : getI4LittleEndian(codeAddr);
                if (compIsForInlining() && jmpDist == 0 && (opcode == CEE_BR || opcode == CEE_BR_S))
                {
                continue; /* NOP */
                }

                So for inlinees only, we will not break blocks because of jump to next. This causes divergence like the following:

                ;; compiling as root
                -----------------------------------------------------------------------------------------------------------------------------------------
                BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
                -----------------------------------------------------------------------------------------------------------------------------------------
                BB01 [0000] 1 1 [000..007)-> BB03 ( cond ) BB02 [0001] 1 BB01 1 [007..016)-> BB04 (always) BB03 [0002] 1 BB01 1 [016..017) BB04 [0003] 2 BB02,BB03 1 [017..039)-> BB05 (always) BB05 [0004] 1 BB04 1 [039..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
                Using edge profiling
                EfficientEdgeCountInstrumentor: preparing for instrumentation
                New BlockSet epoch 1, # of blocks (including unused BB00): 6, bitset array size: 1 (short)
                [0] New probe for BB05 -> BB01 [source]
                [1] New probe for BB03 -> BB04 [source]
                5 blocks, 2 probes (0 on critical edges)
                ;; as inlinee
                -----------------------------------------------------------------------------------------------------------------------------------------
                BBnum BBid ref try hnd preds weight lp [IL range] [jump] [EH region] [flags]
                -----------------------------------------------------------------------------------------------------------------------------------------
                BB01 [0102] 1 100 [000..007)-> BB03 ( cond ) BB02 [0103] 1 BB01 100 [007..016)-> BB04 (always) BB03 [0104] 1 BB01 100 [016..017) BB04 [0105] 2 BB02,BB03 100 [017..06A) (return) -----------------------------------------------------------------------------------------------------------------------------------------
                *************** Inline @[000773] Starting PHASE Profile incorporation
                Have Dynamic PGO: 2 schema records (schema at 000001F14A67CC78, data at 000001F14A63F8E8)
                Profile summary: 1 runs, 0 block probes, 2 edge probes, 0 class profiles, 0 method profiles, 0 other records
                Reconstructing block counts from sparse edge instrumentation
                ... adding known edge BB03 -> BB04: weight 65208
                Could not find source block for schema entry 1 (IL offset/key 00000039)
                ... not solving because of the mismatch
                ... discarding profile count data: PGO data available, but IL did not match
                Computing inlinee profile scale:
                

                We have been relying on the fact that if we build our instrumentation plan super early we always get the same flow graph.

                Seems like the pragmatic thing to do is to always do this optimization if we're optimizing or instrumenting, and not just for inlinees. There is matching logic in the importer and perhaps elsewhere so this should probably be encapsulated into a helper.

                Doing that will (temporarily) break the ability to ingest static PGO for (hopefully just a few) root methods that have branch to next like this. Hopefully not too many of them. And SPMI will have a number of diffs as well, both cases where we no longer read PGO data for root methods and cases where we do read PGO data for inlinees).

                Metadata

                Metadata

                Assignees

                Labels

                area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

                Type

                No type

                Projects

                No projects

                  Milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions