Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Speed3R: Sparse Feed-forward 3D Reconstruction Models

Weining Ren1Xiao Tan2Kai Han1

1The University of Hong Kong 2Baidu AMU

PaperProject PageHugging Face

Speed3R accelerate VGGT and π³ with trainable sparse attention

📣 Updates

  • [April 6, 2026] Training Code Release
  • [March 6, 2026] Initial Release

✨ Overview

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a computational bottleneck that severely limits inference speed.

To resolve this, we introduce Speed3R, an end-to-end trainable model inspired by the core principle of Structure-from-Motion that a sparse set of keypoints is sufficient for robust estimation. Speed3R features a dual-branch attention mechanism where the compression branch creates a coarse contextual prior to guide the selection branch, which performs fine-grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000-view sequences, while introducing a minimal, controlled trade-off in accuracy. Validated on standard benchmarks with both VGGT and π³ backbones, our method delivers high-quality reconstructions at a fraction of computational cost, paving the way for efficient large-scale scene modeling.

🚀 Quick Start

1. Clone & Install Dependencies

First, clone the repository and install the required packages.

git clone https://github.com/Visual-AI/speed3r.git
cd speed3r
pip install -r requirements.txt
pip install triton==3.3.1

2. Run Inference from Command Line

Try our example inference script. You can run it on a directory of images or a video file.

If the automatic download from Hugging Face is slow, you can download the model checkpoint manually from Speed3R_Pi3 and specify its local path using the --ckpt argument.

# Run with the default example video
python example.py

3. Run with Gradio Demo

You can also launch a local Gradio demo for an interactive experience.

# Install demo-specific requirements
pip install -r requirements_demo.txt
# Launch the demo
python demo_gradio.py

🛠️ Detailed Usage

Model Input & Output

The model takes a tensor of images and outputs a dictionary containing the reconstructed geometry.

  • Input: A torch.Tensor of shape $B \times N \times 3 \times H \times W$ with pixel values in the range [0, 1].
  • Output: A dict with the following keys:
    • points: Global point cloud unprojected by local points and camerae_poses (torch.Tensor, $B \times N \times H \times W \times 3$).
    • local_points: Per-view local point maps (torch.Tensor, $B \times N \times H \times W \times 3$).
    • conf: Confidence scores for local points (Raw confidence logits. Apply torch.sigmoid() to obtain probabilities in [0, 1], higher is better) (torch.Tensor, $B \times N \times H \times W \times 1$).
    • camera_poses: Camera-to-world transformation matrices (4x4 in OpenCV format) (torch.Tensor, $B \times N \times 4 \times 4$).

Example Code Snippet

Here is a minimal example of how to run the model on a batch of images.

importtorchfrompi3.models.pi3_sparseimportPi3_Sparsefrompi3.utils.basicimportload_images_as_tensor# Assuming you have a helper function# --- Setup ---device='cuda'iftorch.cuda.is_available() else'cpu'model=Pi3_Sparse.from_pretrained("weining17/Speed3R_Pi3").to(device).eval()
# or download checkpoints from `https://huggingface.co/weining17/Speed3R_Pi3/tree/main/model.safetensors`# --- Load Data ---# Load a sequence of N images into a tensor# imgs shape: (N, 3, H, W).# imgs value: [0, 1]imgs=load_images_as_tensor('path/to/your/data', interval=10).to(device)
# --- Inference ---print("Running model inference...")
# Use mixed precision for better performance on compatible GPUsdtype=torch.bfloat16iftorch.cuda.is_available() andtorch.cuda.get_device_capability()[0] >=8elsetorch.float16withtorch.no_grad():
withtorch.amp.autocast('cuda', dtype=dtype):
# Add a batch dimension -> (1, N, 3, H, W)results=model(imgs[None])
print("Reconstruction complete!")
# Access outputs: results['points'], results['camera_poses'] and results['local_points'].

Training Details

To train the model, please follow VGGT to prepare the CO3Dv2 dataset, and run:

bash scripts/train.sh

TODOs

  • Release the training code
  • Release Speed3R-VGGT code & ckpt

Notice

  1. Currently, the model only supports resolutions that are multiples of 56 rather than 14.
  2. We test the method with triton version 3.3.1, lower version may cause numerical error.
  3. Curently the kernel only support bf16/fp16.

🙏 Acknowledgements

Our work builds upon several fantastic open-source projects. We'd like to express our gratitude to the authors of:

Excellent Concurrent Works Accelerating VGGT

📜 Citation

If you find our work useful, please consider citing:

@article{ren2026speed3r,
title={Speed3R: Sparse Feed-forward 3D Reconstruction Models},
author={Ren, Weining and Tan, Xiao and Han, Kai},
journal={arXiv preprint arXiv:2603.08055},
year={2026}
}

📄 License

This project adopts a dual-licensing strategy following Pi3:

ComponentLicenseCommercial Use
Code (Scripts, Tools, Logic)BSD 3-ClausePermitted
Model Weights (Pi3 Weights)CC BY-NC 4.0Strictly Non-Commercial

Note on Model Weights: Due to the nature of the training datasets, the model weights are restricted to non-commercial research and educational purposes only. Redistribution of the weights must maintain this restriction.

About

[CVPR 2026 Findings] Speed3R: Sparse Feed-forward 3D Reconstruction Models

Resources

Stars

79 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages