Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

PWCPWC

By Chuxin Wang, Wenfei Yang, Xiang Liu, and Tianzhu Zhang.

This repo is the official implementation of "State Space Model Meets Transformer: A New Paradigm for 3D Object Detection".

Introduction

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement.

Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. In this repository, we provide model implementation (with Pytorch) as well as data preparation, training and evaluation scripts on ScanNet and SUN RGB-D.

Interactive State Space Model

In the ISSM, we model the query points as the system states and the scene points as the system inputs. Unlike previous SSMs (Gu et al., 2021a; Gu & Dao, 2023; Dao & Gu, 2024), the proposed ISSM determines how to update the system states based on both the system states and system inputs. Specifically, we modify the SSM parameters (∆, B, C) to be dependent on the system states and design a spatial correlation module to model the relationship between state points and scene points. Therefore, the system states in the ISSM can effectively fulfill the role of queries in complex 3D indoor detection tasks.

Main Results

Visual Comparison on ScanNet V2

Visual Comparison on SUN RGB-D

Usage

  • [Released] Interactive State Space Model, Implementation available at: MultiHeadISSMScan
  • [Released] DEST Framework (based on GroupFree3D), Implementation available at: DestDetector
  • [Planned Releases] Training & Evaluation Pipeline
  • [Planned Releases] DEST Framework (based on VDETR)

Please refer to test_issm.py for using the ISSM module.

Our testing environment is based on python==3.8.19, pytorch==1.12.1+cu113, and triton==3.0.0.

Note that the first execution of the code will be relatively slow due to Triton's JIT (Just-In-Time) compilation process, but subsequent runs will achieve normal speed as the compiled kernels are cached.

Below are our simple test results on a single NVIDIA RTX 3090 GPU:

1: Time taken: 0.0014648370000003297 seconds
2: Time taken: 0.0015144770000006247 seconds
3: Time taken: 0.0014919310000038877 seconds
4: Time taken: 0.001444603999999572 seconds
5: Time taken: 0.001540446000007023 seconds
6: Time taken: 0.0014789960000030078 seconds

Acknowledgements

We thank a lot for the flexible codebase of GroupFree3D and VDETR.

Citation

@inproceedings{wangstate,
title={State Space Model Meets Transformer: A New Paradigm for 3D Object Detection},
author={Wang, Chuxin and Yang, Wenfei and Liu, Xiang and Zhang, Tianzhu},
booktitle={The Thirteenth International Conference on Learning Representations}
}

About

PyTorch implementation for our ICLR 2025 paper State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

Resources

Stars

46 stars

Watchers

3 watching

Forks

Used by

Contributors

Languages