JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

Description

@AndyAyersMS

Background

In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
shift group), @tannergoodingpointed out
that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
duplicate information that's already accurately tracked in the
xarch instruction tables, and risk drifting from reality:

We notably have this info in the instruction tables. For example, see
instrsxarch.h#L1258-L1260,
where ror is tracked as Undefined_OF | Writes_CF which means
other flags (like ZF) are left "as is".

It might be goodness to have most of the trivial arithmetic ops like
this rather query the instruction codegen would emit and lookup the
actual flags, rather than hardcode them all. Particularly because
that may get out of sync, be missed when we fixup the general
tables, or in the future if/when we change the emit strategy for a
given node kind.

Codegen itself typically has such a mapping of Oper -> ins already
and while there is some nuance, the places we typically care about
such queries (lowering) should know statically what codegen will
emit (as this is required for containment and other LSRA required
flags to be correct)

This issue tracks doing that refactor for the xarch path of
SupportsSettingZeroFlag().

Current state (after #129288 fix)

#if defined(TARGET_XARCH)
if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
{
// Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
}
// Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
{
returntrue;
}
#ifdef FEATURE_HW_INTRINSICS
if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
{
returntrue;
}
#endif
#endif

The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
the cleaner pattern.

Building blocks already in tree

  • instrsxarch.h carries accurate per-instruction insFlags
    (Writes_ZF, Undefined_OF, Resets_OF, etc.).
  • emitter::DoesWriteZeroFlag(instruction ins) in
    emitxarch.cpp
    already queries it.
  • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
    codegenxarch.cpp
    already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
    instruction. It uses no instance state, so making it static is
    straightforward.

Proposed shape

#if defined(TARGET_XARCH)
// For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
{
var_types type = TypeGet();
if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
{
instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
{
returnfalse;
}
// Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
{
returngtGetOp2()->IsNeverZero();
}
returntrue;
}
}
#ifdef FEATURE_HW_INTRINSICS
if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
{
returntrue;
}
#endif
#endif

Two variants for getting the Oper -> ins mapping callable from
gentree.cpp:

  • 2a: Refactor genGetInsForOper into a static helper
    (CodeGen::GetIns_FromOper) since it uses no instance state.
  • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
    in instr.cpp covering just the arith opers we care about,
    returning INS_invalid otherwise. Avoids the codegen.h API churn at
    the cost of a small duplicated table.

Note on shift semantics

The instruction tables don't model the Intel SDM "if count is 0 the
flags are unchanged" rule for shl/shr/sar — they show
Writes_ZF unconditionally. The existing
gtGetOp2()->IsNeverZero() guard must be preserved for the shift
case; the table query gates on DoesWriteZeroFlag(ins) being true,
but the count-dependent special case still requires the IsNeverZero
check. This is the only non-trivial nuance in the proposed refactor.

Outcome

  • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
    scalar list (and the shift list) in favor of a single
    table-driven check.
  • Catches future drift automatically when the instruction tables
    change.
  • Aligns the scalar path with the already-existing GT_HWINTRINSIC
    pattern in the same function.

cc @tannergooding@jakobbotsch@EgorBo (last touched
SupportsSettingZeroFlag in #117960)

Metadata

Metadata

Assignees

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
     blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    
    Skip to content

    JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

    Description

    @AndyAyersMS

    Background

    In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
    shift group), @tannergoodingpointed out
    that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
    duplicate information that's already accurately tracked in the
    xarch instruction tables, and risk drifting from reality:

    We notably have this info in the instruction tables. For example, see
    instrsxarch.h#L1258-L1260,
    where ror is tracked as Undefined_OF | Writes_CF which means
    other flags (like ZF) are left "as is".

    It might be goodness to have most of the trivial arithmetic ops like
    this rather query the instruction codegen would emit and lookup the
    actual flags, rather than hardcode them all. Particularly because
    that may get out of sync, be missed when we fixup the general
    tables, or in the future if/when we change the emit strategy for a
    given node kind.

    Codegen itself typically has such a mapping of Oper -> ins already
    and while there is some nuance, the places we typically care about
    such queries (lowering) should know statically what codegen will
    emit (as this is required for containment and other LSRA required
    flags to be correct)

    This issue tracks doing that refactor for the xarch path of
    SupportsSettingZeroFlag().

    Current state (after #129288 fix)

    #if defined(TARGET_XARCH)
    if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
    {
    // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
    }
    // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
    {
    returntrue;
    }
    #ifdef FEATURE_HW_INTRINSICS
    if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
    {
    returntrue;
    }
    #endif
    #endif

    The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
    exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
    the cleaner pattern.

    Building blocks already in tree

    • instrsxarch.h carries accurate per-instruction insFlags
      (Writes_ZF, Undefined_OF, Resets_OF, etc.).
    • emitter::DoesWriteZeroFlag(instruction ins) in
      emitxarch.cpp
      already queries it.
    • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
      codegenxarch.cpp
      already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
      instruction. It uses no instance state, so making it static is
      straightforward.

    Proposed shape

    #if defined(TARGET_XARCH)
    // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
    GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
    {
    var_types type = TypeGet();
    if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
    {
    instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
    if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
    {
    returnfalse;
    }
    // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
    {
    returngtGetOp2()->IsNeverZero();
    }
    returntrue;
    }
    }
    #ifdef FEATURE_HW_INTRINSICS
    if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
    {
    returntrue;
    }
    #endif
    #endif

    Two variants for getting the Oper -> ins mapping callable from
    gentree.cpp:

    • 2a: Refactor genGetInsForOper into a static helper
      (CodeGen::GetIns_FromOper) since it uses no instance state.
    • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
      in instr.cpp covering just the arith opers we care about,
      returning INS_invalid otherwise. Avoids the codegen.h API churn at
      the cost of a small duplicated table.

    Note on shift semantics

    The instruction tables don't model the Intel SDM "if count is 0 the
    flags are unchanged" rule for shl/shr/sar — they show
    Writes_ZF unconditionally. The existing
    gtGetOp2()->IsNeverZero() guard must be preserved for the shift
    case; the table query gates on DoesWriteZeroFlag(ins) being true,
    but the count-dependent special case still requires the IsNeverZero
    check. This is the only non-trivial nuance in the proposed refactor.

    Outcome

    • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
      scalar list (and the shift list) in favor of a single
      table-driven check.
    • Catches future drift automatically when the instruction tables
      change.
    • Aligns the scalar path with the already-existing GT_HWINTRINSIC
      pattern in the same function.

    cc @tannergooding@jakobbotsch@EgorBo (last touched
    SupportsSettingZeroFlag in #117960)

    Metadata

    Metadata

    Assignees

    Labels

    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
      Skip to content

      JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

      Description

      @AndyAyersMS

      Background

      In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
      shift group), @tannergoodingpointed out
      that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
      duplicate information that's already accurately tracked in the
      xarch instruction tables, and risk drifting from reality:

      We notably have this info in the instruction tables. For example, see
      instrsxarch.h#L1258-L1260,
      where ror is tracked as Undefined_OF | Writes_CF which means
      other flags (like ZF) are left "as is".

      It might be goodness to have most of the trivial arithmetic ops like
      this rather query the instruction codegen would emit and lookup the
      actual flags, rather than hardcode them all. Particularly because
      that may get out of sync, be missed when we fixup the general
      tables, or in the future if/when we change the emit strategy for a
      given node kind.

      Codegen itself typically has such a mapping of Oper -> ins already
      and while there is some nuance, the places we typically care about
      such queries (lowering) should know statically what codegen will
      emit (as this is required for containment and other LSRA required
      flags to be correct)

      This issue tracks doing that refactor for the xarch path of
      SupportsSettingZeroFlag().

      Current state (after #129288 fix)

      #if defined(TARGET_XARCH)
      if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
      {
      // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
      }
      // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
      {
      returntrue;
      }
      #ifdef FEATURE_HW_INTRINSICS
      if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
      {
      returntrue;
      }
      #endif
      #endif

      The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
      exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
      the cleaner pattern.

      Building blocks already in tree

      • instrsxarch.h carries accurate per-instruction insFlags
        (Writes_ZF, Undefined_OF, Resets_OF, etc.).
      • emitter::DoesWriteZeroFlag(instruction ins) in
        emitxarch.cpp
        already queries it.
      • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
        codegenxarch.cpp
        already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
        instruction. It uses no instance state, so making it static is
        straightforward.

      Proposed shape

      #if defined(TARGET_XARCH)
      // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
      GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
      {
      var_types type = TypeGet();
      if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
      {
      instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
      if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
      {
      returnfalse;
      }
      // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
      {
      returngtGetOp2()->IsNeverZero();
      }
      returntrue;
      }
      }
      #ifdef FEATURE_HW_INTRINSICS
      if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
      {
      returntrue;
      }
      #endif
      #endif

      Two variants for getting the Oper -> ins mapping callable from
      gentree.cpp:

      • 2a: Refactor genGetInsForOper into a static helper
        (CodeGen::GetIns_FromOper) since it uses no instance state.
      • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
        in instr.cpp covering just the arith opers we care about,
        returning INS_invalid otherwise. Avoids the codegen.h API churn at
        the cost of a small duplicated table.

      Note on shift semantics

      The instruction tables don't model the Intel SDM "if count is 0 the
      flags are unchanged" rule for shl/shr/sar — they show
      Writes_ZF unconditionally. The existing
      gtGetOp2()->IsNeverZero() guard must be preserved for the shift
      case; the table query gates on DoesWriteZeroFlag(ins) being true,
      but the count-dependent special case still requires the IsNeverZero
      check. This is the only non-trivial nuance in the proposed refactor.

      Outcome

      • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
        scalar list (and the shift list) in favor of a single
        table-driven check.
      • Catches future drift automatically when the instruction tables
        change.
      • Aligns the scalar path with the already-existing GT_HWINTRINSIC
        pattern in the same function.

      cc @tannergooding@jakobbotsch@EgorBo (last touched
      SupportsSettingZeroFlag in #117960)

      Metadata

      Metadata

      Assignees

      Labels

      area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

      Type

      No type

      Projects

      No projects

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
        Skip to content

        JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

        Description

        @AndyAyersMS

        Background

        In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
        shift group), @tannergoodingpointed out
        that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
        duplicate information that's already accurately tracked in the
        xarch instruction tables, and risk drifting from reality:

        We notably have this info in the instruction tables. For example, see
        instrsxarch.h#L1258-L1260,
        where ror is tracked as Undefined_OF | Writes_CF which means
        other flags (like ZF) are left "as is".

        It might be goodness to have most of the trivial arithmetic ops like
        this rather query the instruction codegen would emit and lookup the
        actual flags, rather than hardcode them all. Particularly because
        that may get out of sync, be missed when we fixup the general
        tables, or in the future if/when we change the emit strategy for a
        given node kind.

        Codegen itself typically has such a mapping of Oper -> ins already
        and while there is some nuance, the places we typically care about
        such queries (lowering) should know statically what codegen will
        emit (as this is required for containment and other LSRA required
        flags to be correct)

        This issue tracks doing that refactor for the xarch path of
        SupportsSettingZeroFlag().

        Current state (after #129288 fix)

        #if defined(TARGET_XARCH)
        if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
        {
        // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
        }
        // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
        {
        returntrue;
        }
        #ifdef FEATURE_HW_INTRINSICS
        if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
        {
        returntrue;
        }
        #endif
        #endif

        The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
        exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
        the cleaner pattern.

        Building blocks already in tree

        • instrsxarch.h carries accurate per-instruction insFlags
          (Writes_ZF, Undefined_OF, Resets_OF, etc.).
        • emitter::DoesWriteZeroFlag(instruction ins) in
          emitxarch.cpp
          already queries it.
        • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
          codegenxarch.cpp
          already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
          instruction. It uses no instance state, so making it static is
          straightforward.

        Proposed shape

        #if defined(TARGET_XARCH)
        // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
        GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
        {
        var_types type = TypeGet();
        if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
        {
        instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
        if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
        {
        returnfalse;
        }
        // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
        {
        returngtGetOp2()->IsNeverZero();
        }
        returntrue;
        }
        }
        #ifdef FEATURE_HW_INTRINSICS
        if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
        {
        returntrue;
        }
        #endif
        #endif

        Two variants for getting the Oper -> ins mapping callable from
        gentree.cpp:

        • 2a: Refactor genGetInsForOper into a static helper
          (CodeGen::GetIns_FromOper) since it uses no instance state.
        • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
          in instr.cpp covering just the arith opers we care about,
          returning INS_invalid otherwise. Avoids the codegen.h API churn at
          the cost of a small duplicated table.

        Note on shift semantics

        The instruction tables don't model the Intel SDM "if count is 0 the
        flags are unchanged" rule for shl/shr/sar — they show
        Writes_ZF unconditionally. The existing
        gtGetOp2()->IsNeverZero() guard must be preserved for the shift
        case; the table query gates on DoesWriteZeroFlag(ins) being true,
        but the count-dependent special case still requires the IsNeverZero
        check. This is the only non-trivial nuance in the proposed refactor.

        Outcome

        • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
          scalar list (and the shift list) in favor of a single
          table-driven check.
        • Catches future drift automatically when the instruction tables
          change.
        • Aligns the scalar path with the already-existing GT_HWINTRINSIC
          pattern in the same function.

        cc @tannergooding@jakobbotsch@EgorBo (last touched
        SupportsSettingZeroFlag in #117960)

        Metadata

        Metadata

        Assignees

        Labels

        area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

        Type

        No type

        Projects

        No projects

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
          Skip to content

          JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

          Description

          @AndyAyersMS

          Background

          In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
          shift group), @tannergoodingpointed out
          that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
          duplicate information that's already accurately tracked in the
          xarch instruction tables, and risk drifting from reality:

          We notably have this info in the instruction tables. For example, see
          instrsxarch.h#L1258-L1260,
          where ror is tracked as Undefined_OF | Writes_CF which means
          other flags (like ZF) are left "as is".

          It might be goodness to have most of the trivial arithmetic ops like
          this rather query the instruction codegen would emit and lookup the
          actual flags, rather than hardcode them all. Particularly because
          that may get out of sync, be missed when we fixup the general
          tables, or in the future if/when we change the emit strategy for a
          given node kind.

          Codegen itself typically has such a mapping of Oper -> ins already
          and while there is some nuance, the places we typically care about
          such queries (lowering) should know statically what codegen will
          emit (as this is required for containment and other LSRA required
          flags to be correct)

          This issue tracks doing that refactor for the xarch path of
          SupportsSettingZeroFlag().

          Current state (after #129288 fix)

          #if defined(TARGET_XARCH)
          if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
          {
          // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
          }
          // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
          {
          returntrue;
          }
          #ifdef FEATURE_HW_INTRINSICS
          if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
          {
          returntrue;
          }
          #endif
          #endif

          The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
          exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
          the cleaner pattern.

          Building blocks already in tree

          • instrsxarch.h carries accurate per-instruction insFlags
            (Writes_ZF, Undefined_OF, Resets_OF, etc.).
          • emitter::DoesWriteZeroFlag(instruction ins) in
            emitxarch.cpp
            already queries it.
          • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
            codegenxarch.cpp
            already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
            instruction. It uses no instance state, so making it static is
            straightforward.

          Proposed shape

          #if defined(TARGET_XARCH)
          // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
          GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
          {
          var_types type = TypeGet();
          if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
          {
          instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
          if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
          {
          returnfalse;
          }
          // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
          {
          returngtGetOp2()->IsNeverZero();
          }
          returntrue;
          }
          }
          #ifdef FEATURE_HW_INTRINSICS
          if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
          {
          returntrue;
          }
          #endif
          #endif

          Two variants for getting the Oper -> ins mapping callable from
          gentree.cpp:

          • 2a: Refactor genGetInsForOper into a static helper
            (CodeGen::GetIns_FromOper) since it uses no instance state.
          • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
            in instr.cpp covering just the arith opers we care about,
            returning INS_invalid otherwise. Avoids the codegen.h API churn at
            the cost of a small duplicated table.

          Note on shift semantics

          The instruction tables don't model the Intel SDM "if count is 0 the
          flags are unchanged" rule for shl/shr/sar — they show
          Writes_ZF unconditionally. The existing
          gtGetOp2()->IsNeverZero() guard must be preserved for the shift
          case; the table query gates on DoesWriteZeroFlag(ins) being true,
          but the count-dependent special case still requires the IsNeverZero
          check. This is the only non-trivial nuance in the proposed refactor.

          Outcome

          • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
            scalar list (and the shift list) in favor of a single
            table-driven check.
          • Catches future drift automatically when the instruction tables
            change.
          • Aligns the scalar path with the already-existing GT_HWINTRINSIC
            pattern in the same function.

          cc @tannergooding@jakobbotsch@EgorBo (last touched
          SupportsSettingZeroFlag in #117960)

          Metadata

          Metadata

          Assignees

          Labels

          area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

          Type

          No type

          Projects

          No projects

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
            Skip to content

            JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

            Description

            @AndyAyersMS

            Background

            In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
            shift group), @tannergoodingpointed out
            that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
            duplicate information that's already accurately tracked in the
            xarch instruction tables, and risk drifting from reality:

            We notably have this info in the instruction tables. For example, see
            instrsxarch.h#L1258-L1260,
            where ror is tracked as Undefined_OF | Writes_CF which means
            other flags (like ZF) are left "as is".

            It might be goodness to have most of the trivial arithmetic ops like
            this rather query the instruction codegen would emit and lookup the
            actual flags, rather than hardcode them all. Particularly because
            that may get out of sync, be missed when we fixup the general
            tables, or in the future if/when we change the emit strategy for a
            given node kind.

            Codegen itself typically has such a mapping of Oper -> ins already
            and while there is some nuance, the places we typically care about
            such queries (lowering) should know statically what codegen will
            emit (as this is required for containment and other LSRA required
            flags to be correct)

            This issue tracks doing that refactor for the xarch path of
            SupportsSettingZeroFlag().

            Current state (after #129288 fix)

            #if defined(TARGET_XARCH)
            if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
            {
            // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
            }
            // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
            {
            returntrue;
            }
            #ifdef FEATURE_HW_INTRINSICS
            if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
            {
            returntrue;
            }
            #endif
            #endif

            The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
            exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
            the cleaner pattern.

            Building blocks already in tree

            • instrsxarch.h carries accurate per-instruction insFlags
              (Writes_ZF, Undefined_OF, Resets_OF, etc.).
            • emitter::DoesWriteZeroFlag(instruction ins) in
              emitxarch.cpp
              already queries it.
            • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
              codegenxarch.cpp
              already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
              instruction. It uses no instance state, so making it static is
              straightforward.

            Proposed shape

            #if defined(TARGET_XARCH)
            // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
            GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
            {
            var_types type = TypeGet();
            if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
            {
            instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
            if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
            {
            returnfalse;
            }
            // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
            {
            returngtGetOp2()->IsNeverZero();
            }
            returntrue;
            }
            }
            #ifdef FEATURE_HW_INTRINSICS
            if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
            {
            returntrue;
            }
            #endif
            #endif

            Two variants for getting the Oper -> ins mapping callable from
            gentree.cpp:

            • 2a: Refactor genGetInsForOper into a static helper
              (CodeGen::GetIns_FromOper) since it uses no instance state.
            • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
              in instr.cpp covering just the arith opers we care about,
              returning INS_invalid otherwise. Avoids the codegen.h API churn at
              the cost of a small duplicated table.

            Note on shift semantics

            The instruction tables don't model the Intel SDM "if count is 0 the
            flags are unchanged" rule for shl/shr/sar — they show
            Writes_ZF unconditionally. The existing
            gtGetOp2()->IsNeverZero() guard must be preserved for the shift
            case; the table query gates on DoesWriteZeroFlag(ins) being true,
            but the count-dependent special case still requires the IsNeverZero
            check. This is the only non-trivial nuance in the proposed refactor.

            Outcome

            • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
              scalar list (and the shift list) in favor of a single
              table-driven check.
            • Catches future drift automatically when the instruction tables
              change.
            • Aligns the scalar path with the already-existing GT_HWINTRINSIC
              pattern in the same function.

            cc @tannergooding@jakobbotsch@EgorBo (last touched
            SupportsSettingZeroFlag in #117960)

            Metadata

            Metadata

            Assignees

            Labels

            area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

            Type

            No type

            Projects

            No projects

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

              Description

              @AndyAyersMS

              Background

              In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
              shift group), @tannergoodingpointed out
              that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
              duplicate information that's already accurately tracked in the
              xarch instruction tables, and risk drifting from reality:

              We notably have this info in the instruction tables. For example, see
              instrsxarch.h#L1258-L1260,
              where ror is tracked as Undefined_OF | Writes_CF which means
              other flags (like ZF) are left "as is".

              It might be goodness to have most of the trivial arithmetic ops like
              this rather query the instruction codegen would emit and lookup the
              actual flags, rather than hardcode them all. Particularly because
              that may get out of sync, be missed when we fixup the general
              tables, or in the future if/when we change the emit strategy for a
              given node kind.

              Codegen itself typically has such a mapping of Oper -> ins already
              and while there is some nuance, the places we typically care about
              such queries (lowering) should know statically what codegen will
              emit (as this is required for containment and other LSRA required
              flags to be correct)

              This issue tracks doing that refactor for the xarch path of
              SupportsSettingZeroFlag().

              Current state (after #129288 fix)

              #if defined(TARGET_XARCH)
              if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
              {
              // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
              }
              // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
              {
              returntrue;
              }
              #ifdef FEATURE_HW_INTRINSICS
              if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
              {
              returntrue;
              }
              #endif
              #endif

              The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
              exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
              the cleaner pattern.

              Building blocks already in tree

              • instrsxarch.h carries accurate per-instruction insFlags
                (Writes_ZF, Undefined_OF, Resets_OF, etc.).
              • emitter::DoesWriteZeroFlag(instruction ins) in
                emitxarch.cpp
                already queries it.
              • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
                codegenxarch.cpp
                already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
                instruction. It uses no instance state, so making it static is
                straightforward.

              Proposed shape

              #if defined(TARGET_XARCH)
              // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
              GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
              {
              var_types type = TypeGet();
              if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
              {
              instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
              if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
              {
              returnfalse;
              }
              // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
              {
              returngtGetOp2()->IsNeverZero();
              }
              returntrue;
              }
              }
              #ifdef FEATURE_HW_INTRINSICS
              if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
              {
              returntrue;
              }
              #endif
              #endif

              Two variants for getting the Oper -> ins mapping callable from
              gentree.cpp:

              • 2a: Refactor genGetInsForOper into a static helper
                (CodeGen::GetIns_FromOper) since it uses no instance state.
              • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
                in instr.cpp covering just the arith opers we care about,
                returning INS_invalid otherwise. Avoids the codegen.h API churn at
                the cost of a small duplicated table.

              Note on shift semantics

              The instruction tables don't model the Intel SDM "if count is 0 the
              flags are unchanged" rule for shl/shr/sar — they show
              Writes_ZF unconditionally. The existing
              gtGetOp2()->IsNeverZero() guard must be preserved for the shift
              case; the table query gates on DoesWriteZeroFlag(ins) being true,
              but the count-dependent special case still requires the IsNeverZero
              check. This is the only non-trivial nuance in the proposed refactor.

              Outcome

              • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
                scalar list (and the shift list) in favor of a single
                table-driven check.
              • Catches future drift automatically when the instruction tables
                change.
              • Aligns the scalar path with the already-existing GT_HWINTRINSIC
                pattern in the same function.

              cc @tannergooding@jakobbotsch@EgorBo (last touched
              SupportsSettingZeroFlag in #117960)

              Metadata

              Metadata

              Assignees

              Labels

              area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

              Type

              No type

              Projects

              No projects

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                Skip to content

                JIT: SupportsSettingZeroFlag xarch path should query instruction tables instead of hardcoding op lists #129346

                Description

                @AndyAyersMS

                Background

                In #129288 (fixed by the minimal removal of GT_ROL/GT_ROR from the
                shift group), @tannergoodingpointed out
                that hardcoded op→flag lists in GenTree::SupportsSettingZeroFlag()
                duplicate information that's already accurately tracked in the
                xarch instruction tables, and risk drifting from reality:

                We notably have this info in the instruction tables. For example, see
                instrsxarch.h#L1258-L1260,
                where ror is tracked as Undefined_OF | Writes_CF which means
                other flags (like ZF) are left "as is".

                It might be goodness to have most of the trivial arithmetic ops like
                this rather query the instruction codegen would emit and lookup the
                actual flags, rather than hardcode them all. Particularly because
                that may get out of sync, be missed when we fixup the general
                tables, or in the future if/when we change the emit strategy for a
                given node kind.

                Codegen itself typically has such a mapping of Oper -> ins already
                and while there is some nuance, the places we typically care about
                such queries (lowering) should know statically what codegen will
                emit (as this is required for containment and other LSRA required
                flags to be correct)

                This issue tracks doing that refactor for the xarch path of
                SupportsSettingZeroFlag().

                Current state (after #129288 fix)

                #if defined(TARGET_XARCH)
                if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
                {
                // Shift instructions do not update the flags in case of count being zero.returngtGetOp2()->IsNeverZero();
                }
                // Note: GT_ROL / GT_ROR are intentionally NOT in this list. On xarch,// ROL/ROR/RCL/RCR only update CF (and OF in the 1-bit form); SF/ZF/AF/PF// are not affected for any count.if (OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG))
                {
                returntrue;
                }
                #ifdef FEATURE_HW_INTRINSICS
                if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
                {
                returntrue;
                }
                #endif
                #endif

                The hardcoded scalar list (GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB, GT_NEG) and the separate shift list still mirror data that already
                exists in instrsxarch.h. The GT_HWINTRINSIC branch already shows
                the cleaner pattern.

                Building blocks already in tree

                • instrsxarch.h carries accurate per-instruction insFlags
                  (Writes_ZF, Undefined_OF, Resets_OF, etc.).
                • emitter::DoesWriteZeroFlag(instruction ins) in
                  emitxarch.cpp
                  already queries it.
                • CodeGen::genGetInsForOper(genTreeOps oper, var_types type) in
                  codegenxarch.cpp
                  already maps GT_ADD, GT_AND, GT_LSH, GT_MUL, GT_NEG, GT_NOT, GT_OR, GT_ROL, GT_ROR, GT_RSH, GT_RSZ, GT_SUB, GT_XOR, ... to the right
                  instruction. It uses no instance state, so making it static is
                  straightforward.

                Proposed shape

                #if defined(TARGET_XARCH)
                // For arith opers whose lowering emits a single primary instruction// we can name statically, defer to the instruction tables.if (OperIs(GT_ADD, GT_AND, GT_OR, GT_XOR, GT_SUB, GT_NEG,
                GT_LSH, GT_RSH, GT_RSZ, GT_ROL, GT_ROR))
                {
                var_types type = TypeGet();
                if (!varTypeIsFloating(type) && !varTypeIsSIMD(type))
                {
                instruction ins = CodeGen::GetIns_FromOper(OperGet(), type);
                if (ins != INS_invalid && !emitter::DoesWriteZeroFlag(ins))
                {
                returnfalse;
                }
                // Shifts only write flags when count != 0 (Intel SDM).// The flag tables can't express the count-dependent// condition; encode it here.if (OperIs(GT_LSH, GT_RSH, GT_RSZ))
                {
                returngtGetOp2()->IsNeverZero();
                }
                returntrue;
                }
                }
                #ifdef FEATURE_HW_INTRINSICS
                if (OperIs(GT_HWINTRINSIC) && emitter::DoesWriteZeroFlag(HWIntrinsicInfo::lookupIns(AsHWIntrinsic(), nullptr)))
                {
                returntrue;
                }
                #endif
                #endif

                Two variants for getting the Oper -> ins mapping callable from
                gentree.cpp:

                • 2a: Refactor genGetInsForOper into a static helper
                  (CodeGen::GetIns_FromOper) since it uses no instance state.
                • 2b: Add a small local GetXArchArithIns(genTreeOps, var_types)
                  in instr.cpp covering just the arith opers we care about,
                  returning INS_invalid otherwise. Avoids the codegen.h API churn at
                  the cost of a small duplicated table.

                Note on shift semantics

                The instruction tables don't model the Intel SDM "if count is 0 the
                flags are unchanged" rule for shl/shr/sar — they show
                Writes_ZF unconditionally. The existing
                gtGetOp2()->IsNeverZero() guard must be preserved for the shift
                case; the table query gates on DoesWriteZeroFlag(ins) being true,
                but the count-dependent special case still requires the IsNeverZero
                check. This is the only non-trivial nuance in the proposed refactor.

                Outcome

                • Removes the hardcoded GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB/GT_NEG
                  scalar list (and the shift list) in favor of a single
                  table-driven check.
                • Catches future drift automatically when the instruction tables
                  change.
                • Aligns the scalar path with the already-existing GT_HWINTRINSIC
                  pattern in the same function.

                cc @tannergooding@jakobbotsch@EgorBo (last touched
                SupportsSettingZeroFlag in #117960)

                Metadata

                Metadata

                Assignees

                Labels

                area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

                Type

                No type

                Projects

                No projects

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions