Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

image

Aether: Geometric-Aware Unified World Modeling

Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling for human-like spatial reasoning. Our framework unifies three core capabilities: (1) 🌏 4D dynamic reconstruction, (2) 🎬 action-conditioned video prediction, and (3) 🎯 goal-conditioned visual planning. Trained entirely on synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.

Teaser

🥳 NEWS:

  • Oct.22nd 2025: Aether won the Outstanding Paper Award at the ICCV 2025 RIWM workshop!
  • Jun.26th 2025: Aether is accepted by ICCV 2025!
  • Jun.3rd 2025:DeepVerse is released! It is a 4D auto-regressive world model. Check it out!
  • Mar.31st 2025: The Gradio demo is available! You can deploy locally or experience Aether online on Hugging Face.
  • Mar.28th 2025: AetherV1 is released! Model checkpoints, paper, website, and inference code are all available.

🔨 Installation

Note: We recommend using virtual environments such as Anaconda.

# clone projectgit clone https://github.com/OpenRobotLab/Aether.gitcd Aether
# create conda environmentconda create -n aether python=3.10conda activate aether
# install dependenciespip install -r requirements.txt

🚀 Inference

Warning: When doing reconstruction, Aether pipeline automatically centers crop the input video if its size does not match 480x720. Therefore, for evaluation purpose, we have to slide a 480p window both on the spatial and temporal dimensions, and blend all windows' outputs both spatially and temporally. Examples of video depth and camera pose evaluation can be found at evaluation/.

Run inference demo locally

  • 4D reconstruction:

    python scripts/demo.py --task reconstruction --video ./assets/example_videos/moviegen.mp4
  • Action-conditioned video prediction:

    python scripts/demo.py --task prediction --image ./assets/example_obs/car.png --raymap_action assets/example_raymaps/raymap_forward_right.npy
  • Goal-conditioned visual planning:

    python scripts/demo.py --task planning --image ./assets/example_obs_goal/01_obs.png --goal ./assets/example_obs_goal/01_goal.png

Results will be saved in ./outputs/ by default.

Run inference demo with Gradio

The Gradio demo provides an interactive web-based Aether experience.

python scripts/demo_gradio.py

Our local testing environment is deployed using an A100 GPU with 80GB of memory, and it is set to run on the local port 7860 by default.

Inference with your own raymap action

Suppose you have a sequence of camera poses, you have to convert it to raymap action trajectories before inference with Aether. Note that your camera poses should be within the camera coordinate system of the first frame. You can use the camera_pose_to_raymap function in postprocess_utils.py.

# suppose you have the ground-truth depth values:disparity=1./depth[depth>0]
dmax=disparity.max()
# otherwise, you can set dmax to 1.0 by default:dmax=1.0# then suppose we have a camera trajectory# camera_pose: shape of (N, 4, 4), e.g. N = 41# intrinsic: shape of (N, 3, 3), e.g. N = 41fromaether.utils.postprocess_utilsimportcamera_pose_to_raymap# we will get a raymap sequence of shape (N, 6, h, w) # where h = image height // 8 and w = image width // 8raymap=camera_pose_to_raymap(camera_pose=camera_pose, intrinsic=intrinsic, dmax=dmax)
# save the raymapnp.save("/path/to/your/raymap.npy", raymap)

📝 Citation

If you find this work useful in your research, please consider citing:

@article{aether,
title = {Aether: Geometric-Aware Unified World Modeling},
author = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
journal = {arXiv preprint arXiv:2503.18945},
year = {2025}
}

💡 Limitations

Aether represents an initial step in our journey, trained entirely on synthetic data. While it demonstrates promising capabilities, it is important to be aware of its current limitations:

  • 🔄 Aether struggles with highly dynamic scenarios, such as those involving significant motion or dense crowds.
  • 📸 Its camera pose estimation can be less stable in certain conditions.
  • 📐 For visual planning tasks, we recommend keeping the observations and goals relatively close to ensure optimal performance.

We are actively working on the next generation of Aether and are committed to addressing these limitations in future releases.

📚 License

This repository is licensed under the MIT License - see the LICENSE file for details. For any questions, please email to tonghe90[at]gmail[dot]com.

✨ Acknowledgements

Our work is primarily built upon Accelerate, Diffusers, CogVideoX, Finetrainers, DepthAnyVideo, CUT3R, MonST3R, VBench, GST, SPA, DroidCalib, Grounded-SAM-2, ceres-solver, etc. We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

Releases

Packages

Used by

Contributors

Languages