Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

⚠️ Note: This repository currently contains only the initial code. Full code and pretrained weights will be updated soon.

[CVPR 2026] 🌀Diff4Splat

Diff4Splat is a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets.

Here is our Project Page.

Feel free to contact us or open an issue if you have any questions or suggestions.

🔥 See Also

You may also be interested in our other works:

  • [CVPR 2026] MoVieS: a feed-forward model for 4D dynamic reconstruction from monocular videos.

📢 News

  • 2026-02-21: The paper is accepted to CVPR 2026.
  • 2025-11-01: Diff4Splat is released on arXiv.
  • 2025-10-15: Initial codebase structure established.
  • 2025-10-01: Project development started.

📝 Paper vs Released Code

This section clarifies the differences between the paper description and the current released codebase, to help set expectations for reproducibility.

What the Paper Describes

  • Video Backbone: CogVideoX-style Video DiT with 32-channel 3D Causal VAE (4×8×8 compression)
  • LDRM Input: Video latent tensor (z) from the diffusion model, together with camera information, processed by LDRM Transformer to output deformable 3D Gaussians

Release Roadmap

We are actively working on:

  1. Paper-faithful implementation: A version closer to the CogVideoX + latent-input LDRM stack described in the paper
  2. Complete training/inference scripts: Exact scripts to reproduce the paper's results
  3. Pretrained checkpoints: Both the paper's setup and this repository's engineering variant

📋 Project Status

  • Inference code released
  • Training code and data preprocessing scripts released
  • Pretrained checkpoints (coming soon)
  • HuggingFace demo (coming soon)
  • Preprocessed dataset (coming soon)

🔧 Installation

Requirements

  • Python >= 3.10
  • PyTorch >= 2.0 (with CUDA support)
  • CUDA >= 11.8

Install Dependencies

# Clone the repository
git clone https://github.com/paulpanwang/Diff4Splat.git
cd Diff4Splat
# Install required packages
pip install -r settings/requirements.txt

📊 Datasets

Configure your dataset root path in src/options.py or via the DATASET_ROOT environment variable.

The following datasets are supported:

  • RealEstate10K (re10k) - Static scenes
  • TartanAir (tartanair) - Static scenes
  • MatrixCity (matrixcity) - Static scenes
  • DL3DV (dl3dv) - Static scenes
  • DynamicReplica (dynamicreplica) - Dynamic scenes
  • PointOdyssey (pointodyssey) - Dynamic scenes
  • VKITTI2 (vkitti2) - Dynamic scenes
  • Spring (spring) - Dynamic scenes
  • Stereo4D (stereo4d) - Dynamic scenes

Dataset paths can be configured in src/options.py:

Questions & Feedback

If you have questions about reproducibility or comparisons, please open an issue or contact the authors. We appreciate your understanding as we continue to improve and complete this codebase!

📚 Citation

If you find our work helpful, please consider citing:

@article{pan2025diff4splat,
title={Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models},
author={Pan, Panwang and Lin, Chenguo and Zhao, Jingjing and Li, Chenxin and Lin, Yuchen and Li, Haopeng and Yan, Honglei and Wen, Kairun and Lin, Yunlong and Yuan, Yixuan and Mu, Yadong},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year={2026}
}

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

😊 Acknowledgement

We would like to thank the authors of MoVieS, PartCrafter, DiffSplat, and other related works for their inspiring research and open-source contributions that helped shape this project.

About

[CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models.

Topics

Resources

Contributing

Stars

121 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages