Skip to content

[Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

Description

@lazycal

When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
as shown in the following code snippet:

importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
defbefore():
x=relay.var("x", shape=x_shape)
weight=relay.var("weight", shape=w_shape)
y=relay.nn.conv2d(
x,
weight,
kernel_size=(3, 3),
padding=(1, 1),
data_layout="NCHW4c",
kernel_layout="OIHW4i4o",
)
y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
y=relay.Function([x, weight], y)
returntvm.IRModule.from_expr(y)
defalter_conv2d(attrs, inputs, tinfos, out_type):
data, weight=inputsnew_attrs=dict(attrs)
new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
be=transform.InferType()(before())
print('='*40, 'before', '='*40)
print(be)
af=transform.AlterOpLayout()(be)
print('='*40, 'after', '='*40)
print(af)
xnp=np.random.rand(*x_shape).astype(np.float32)
wnp=np.random.rand(*w_shape).astype(np.float32)
be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

The module before:

def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
%0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
}

and after:

def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
%0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
%1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
%2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
%3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
}

Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

It seems that StridedSliceInferCorrectLayout is responsible for this.

BTW, the layout_transform seems weird in the latter IR:

%3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

Environment

  • TVM: commit e334942
  • CUDA version: 10.0
  • System: Ubuntu 16.04
  • GCC 5.4
  • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
    Skip to content

    [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

    Description

    @lazycal

    When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
    as shown in the following code snippet:

    importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
    defbefore():
    x=relay.var("x", shape=x_shape)
    weight=relay.var("weight", shape=w_shape)
    y=relay.nn.conv2d(
    x,
    weight,
    kernel_size=(3, 3),
    padding=(1, 1),
    data_layout="NCHW4c",
    kernel_layout="OIHW4i4o",
    )
    y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
    y=relay.Function([x, weight], y)
    returntvm.IRModule.from_expr(y)
    defalter_conv2d(attrs, inputs, tinfos, out_type):
    data, weight=inputsnew_attrs=dict(attrs)
    new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
    withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
    be=transform.InferType()(before())
    print('='*40, 'before', '='*40)
    print(be)
    af=transform.AlterOpLayout()(be)
    print('='*40, 'after', '='*40)
    print(af)
    xnp=np.random.rand(*x_shape).astype(np.float32)
    wnp=np.random.rand(*w_shape).astype(np.float32)
    be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
    af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
    tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
    test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

    The module before:

    def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
    %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
    strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
    }

    and after:

    def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
    %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
    %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
    %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
    %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
    layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
    }

    Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

    It seems that StridedSliceInferCorrectLayout is responsible for this.

    BTW, the layout_transform seems weird in the latter IR:

    %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
    layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

    The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

    Environment

    • TVM: commit e334942
    • CUDA version: 10.0
    • System: Ubuntu 16.04
    • GCC 5.4
    • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

    Metadata

    Metadata

    Assignees

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
      Skip to content

      [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

      Description

      @lazycal

      When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
      as shown in the following code snippet:

      importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
      defbefore():
      x=relay.var("x", shape=x_shape)
      weight=relay.var("weight", shape=w_shape)
      y=relay.nn.conv2d(
      x,
      weight,
      kernel_size=(3, 3),
      padding=(1, 1),
      data_layout="NCHW4c",
      kernel_layout="OIHW4i4o",
      )
      y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
      y=relay.Function([x, weight], y)
      returntvm.IRModule.from_expr(y)
      defalter_conv2d(attrs, inputs, tinfos, out_type):
      data, weight=inputsnew_attrs=dict(attrs)
      new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
      withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
      be=transform.InferType()(before())
      print('='*40, 'before', '='*40)
      print(be)
      af=transform.AlterOpLayout()(be)
      print('='*40, 'after', '='*40)
      print(af)
      xnp=np.random.rand(*x_shape).astype(np.float32)
      wnp=np.random.rand(*w_shape).astype(np.float32)
      be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
      af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
      tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
      test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

      The module before:

      def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
      %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
      strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
      }

      and after:

      def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
      %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
      %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
      %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
      %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
      layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
      }

      Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

      It seems that StridedSliceInferCorrectLayout is responsible for this.

      BTW, the layout_transform seems weird in the latter IR:

      %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
      layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

      The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

      Environment

      • TVM: commit e334942
      • CUDA version: 10.0
      • System: Ubuntu 16.04
      • GCC 5.4
      • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

      Metadata

      Metadata

      Assignees

      Labels

      No labels
      No labels

      Type

      No type

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
        Skip to content

        [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

        Description

        @lazycal

        When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
        as shown in the following code snippet:

        importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
        defbefore():
        x=relay.var("x", shape=x_shape)
        weight=relay.var("weight", shape=w_shape)
        y=relay.nn.conv2d(
        x,
        weight,
        kernel_size=(3, 3),
        padding=(1, 1),
        data_layout="NCHW4c",
        kernel_layout="OIHW4i4o",
        )
        y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
        y=relay.Function([x, weight], y)
        returntvm.IRModule.from_expr(y)
        defalter_conv2d(attrs, inputs, tinfos, out_type):
        data, weight=inputsnew_attrs=dict(attrs)
        new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
        withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
        be=transform.InferType()(before())
        print('='*40, 'before', '='*40)
        print(be)
        af=transform.AlterOpLayout()(be)
        print('='*40, 'after', '='*40)
        print(af)
        xnp=np.random.rand(*x_shape).astype(np.float32)
        wnp=np.random.rand(*w_shape).astype(np.float32)
        be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
        af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
        tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
        test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

        The module before:

        def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
        %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
        strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
        }

        and after:

        def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
        %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
        %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
        %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
        %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
        layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
        }

        Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

        It seems that StridedSliceInferCorrectLayout is responsible for this.

        BTW, the layout_transform seems weird in the latter IR:

        %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
        layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

        The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

        Environment

        • TVM: commit e334942
        • CUDA version: 10.0
        • System: Ubuntu 16.04
        • GCC 5.4
        • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

        Metadata

        Metadata

        Assignees

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
          Skip to content

          [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

          Description

          @lazycal

          When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
          as shown in the following code snippet:

          importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
          defbefore():
          x=relay.var("x", shape=x_shape)
          weight=relay.var("weight", shape=w_shape)
          y=relay.nn.conv2d(
          x,
          weight,
          kernel_size=(3, 3),
          padding=(1, 1),
          data_layout="NCHW4c",
          kernel_layout="OIHW4i4o",
          )
          y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
          y=relay.Function([x, weight], y)
          returntvm.IRModule.from_expr(y)
          defalter_conv2d(attrs, inputs, tinfos, out_type):
          data, weight=inputsnew_attrs=dict(attrs)
          new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
          withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
          be=transform.InferType()(before())
          print('='*40, 'before', '='*40)
          print(be)
          af=transform.AlterOpLayout()(be)
          print('='*40, 'after', '='*40)
          print(af)
          xnp=np.random.rand(*x_shape).astype(np.float32)
          wnp=np.random.rand(*w_shape).astype(np.float32)
          be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
          af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
          tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
          test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

          The module before:

          def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
          %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
          strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
          }

          and after:

          def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
          %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
          %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
          %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
          %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
          layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
          }

          Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

          It seems that StridedSliceInferCorrectLayout is responsible for this.

          BTW, the layout_transform seems weird in the latter IR:

          %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
          layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

          The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

          Environment

          • TVM: commit e334942
          • CUDA version: 10.0
          • System: Ubuntu 16.04
          • GCC 5.4
          • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

          Metadata

          Metadata

          Assignees

          Labels

          No labels
          No labels

          Type

          No type

          Projects

          No projects

            Milestone

            No milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
            Skip to content

            [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

            Description

            @lazycal

            When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
            as shown in the following code snippet:

            importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
            defbefore():
            x=relay.var("x", shape=x_shape)
            weight=relay.var("weight", shape=w_shape)
            y=relay.nn.conv2d(
            x,
            weight,
            kernel_size=(3, 3),
            padding=(1, 1),
            data_layout="NCHW4c",
            kernel_layout="OIHW4i4o",
            )
            y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
            y=relay.Function([x, weight], y)
            returntvm.IRModule.from_expr(y)
            defalter_conv2d(attrs, inputs, tinfos, out_type):
            data, weight=inputsnew_attrs=dict(attrs)
            new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
            withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
            be=transform.InferType()(before())
            print('='*40, 'before', '='*40)
            print(be)
            af=transform.AlterOpLayout()(be)
            print('='*40, 'after', '='*40)
            print(af)
            xnp=np.random.rand(*x_shape).astype(np.float32)
            wnp=np.random.rand(*w_shape).astype(np.float32)
            be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
            af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
            tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
            test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

            The module before:

            def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
            %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
            strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
            }

            and after:

            def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
            %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
            %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
            %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
            %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
            layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
            }

            Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

            It seems that StridedSliceInferCorrectLayout is responsible for this.

            BTW, the layout_transform seems weird in the latter IR:

            %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
            layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

            The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

            Environment

            • TVM: commit e334942
            • CUDA version: 10.0
            • System: Ubuntu 16.04
            • GCC 5.4
            • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

            Metadata

            Metadata

            Assignees

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
              Skip to content

              [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

              Description

              @lazycal

              When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
              as shown in the following code snippet:

              importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
              defbefore():
              x=relay.var("x", shape=x_shape)
              weight=relay.var("weight", shape=w_shape)
              y=relay.nn.conv2d(
              x,
              weight,
              kernel_size=(3, 3),
              padding=(1, 1),
              data_layout="NCHW4c",
              kernel_layout="OIHW4i4o",
              )
              y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
              y=relay.Function([x, weight], y)
              returntvm.IRModule.from_expr(y)
              defalter_conv2d(attrs, inputs, tinfos, out_type):
              data, weight=inputsnew_attrs=dict(attrs)
              new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
              withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
              be=transform.InferType()(before())
              print('='*40, 'before', '='*40)
              print(be)
              af=transform.AlterOpLayout()(be)
              print('='*40, 'after', '='*40)
              print(af)
              xnp=np.random.rand(*x_shape).astype(np.float32)
              wnp=np.random.rand(*w_shape).astype(np.float32)
              be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
              af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
              tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
              test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

              The module before:

              def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
              %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
              strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
              }

              and after:

              def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
              %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
              %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
              %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
              %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
              layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
              }

              Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

              It seems that StridedSliceInferCorrectLayout is responsible for this.

              BTW, the layout_transform seems weird in the latter IR:

              %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
              layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

              The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

              Environment

              • TVM: commit e334942
              • CUDA version: 10.0
              • System: Ubuntu 16.04
              • GCC 5.4
              • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

              Metadata

              Metadata

              Assignees

              Labels

              No labels
              No labels

              Type

              No type

              Projects

              No projects

                Milestone

                No milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms · Issue #8759 · apache/tvm · GitHub
                Skip to content

                [Bug] AlterLayout doesn't correctly wrap strided_slice with layout_transforms #8759

                Description

                @lazycal

                When doing AlterLayout pass on a conv followed by a strided_slice from NCHW4c to NCHW, the compiler does nothing to strided_slice, while (I think) the only correct behavior should be wrapping it with two layout_transforms. This leads to incorrect numerical result/crash at InferType, depending on the concrete input shape,
                as shown in the following code snippet:

                importtvmfromtvmimportrelayfromtvm.relayimporttransformfromtvm.relay.testing.temp_op_attrimportTempOpAttrimportnumpyasnpdeftest1(x_shape, w_shape):
                defbefore():
                x=relay.var("x", shape=x_shape)
                weight=relay.var("weight", shape=w_shape)
                y=relay.nn.conv2d(
                x,
                weight,
                kernel_size=(3, 3),
                padding=(1, 1),
                data_layout="NCHW4c",
                kernel_layout="OIHW4i4o",
                )
                y=relay.strided_slice(y, begin=[0, 0], end=[1, -1], strides=[1, 8])
                y=relay.Function([x, weight], y)
                returntvm.IRModule.from_expr(y)
                defalter_conv2d(attrs, inputs, tinfos, out_type):
                data, weight=inputsnew_attrs=dict(attrs)
                new_attrs["data_layout"] ="NCHW"new_attrs["kernel_layout"] ="OIHW"returnrelay.nn.conv2d(data, weight, **new_attrs)
                withTempOpAttr("nn.conv2d", "FTVMAlterOpLayout", alter_conv2d):
                be=transform.InferType()(before())
                print('='*40, 'before', '='*40)
                print(be)
                af=transform.AlterOpLayout()(be)
                print('='*40, 'after', '='*40)
                print(af)
                xnp=np.random.rand(*x_shape).astype(np.float32)
                wnp=np.random.rand(*w_shape).astype(np.float32)
                be_res=relay.create_executor("debug", be).evaluate()(xnp, wnp).numpy()
                af_res=relay.create_executor("debug", af).evaluate()(xnp, wnp).numpy()
                tvm.testing.assert_allclose(be_res, af_res, rtol=1e-3, atol=1e-3)
                test1(x_shape=(1, 1, 1, 1, 4), w_shape=(9, 1, 3, 3, 4, 4)) # incorrect numerical result# test1(x_shape=(1, 1, 1, 1, 4), w_shape=(11, 1, 3, 3, 4, 4)) # crash at InferType

                The module before:

                def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
                %0 = nn.conv2d(%x, %weight, padding=[1, 1, 1, 1], kernel_size=[3, 3], data_layout="NCHW4c", kernel_layout="OIHW4i4o") /* ty=Tensor[(1, 9, 1, 1, 4), float32] */;
                strided_slice(%0, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
                }

                and after:

                def @main(%x: Tensor[(1, 1, 1, 1, 4), float32], %weight: Tensor[(9, 1, 3, 3, 4, 4), float32]) -> Tensor[(1, 1, 1, 1, 4), float32] {
                %0 = layout_transform(%x, src_layout="NCHW4c", dst_layout="NCHW") /* ty=Tensor[(1, 4, 1, 1), float32] */;
                %1 = layout_transform(%weight, src_layout="OIHW4i4o", dst_layout="OIHW") /* ty=Tensor[(36, 4, 3, 3), float32] */;
                %2 = nn.conv2d(%0, %1, padding=[1, 1, 1, 1], kernel_size=[3, 3]) /* ty=Tensor[(1, 36, 1, 1), float32] */;
                %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
                layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */
                }

                Specifically, I am doing conv_NCHW4c_out[;,::8,...] (a 8-stride slice at the primal C dimension of NCHW4c). After altering layout into NCHW, the compiler does not wrap strided_slice with any layout_transformations nor adjust its attributes, so the semantic gets changed to conv_NCHW_out[:,::8,...], which means picking 1 every 8 elements, while what we need is to pick 4 elements every 4*8=32 elements for conv_NCHW_out

                It seems that StridedSliceInferCorrectLayout is responsible for this.

                BTW, the layout_transform seems weird in the latter IR:

                %3 = strided_slice(%2, begin=[0, 0], end=[1, -1], strides=[1, 8], axes=None) /* ty=Tensor[(1, 5, 1, 1), float32] */;
                layout_transform(%3, src_layout="NCHW", dst_layout="NCHW4c") /* ty=Tensor[(1, 1, 1, 1, 4), float32] */

                The resultant tensor has smaller shape (1,1,1,1,4) than the before-transform one (1,5,1,1), and the reason I think is that (1,5,1,1) is not a valid input to be converted to the layout of NCHW4c, and I thought layout_transform should be able to detect and reject that?

                Environment

                • TVM: commit e334942
                • CUDA version: 10.0
                • System: Ubuntu 16.04
                • GCC 5.4
                • Build options: -DUSE_RELAY_DEBUG=ON -DUSE_CUBLAS=ON -DUSE_LLVM=ON -DUSE_CUDA=ON -DUSE_CUDNN=ON

                Metadata

                Metadata

                Assignees

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions