Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry [CVPR 2026]

DemoModelProject PagePaper

LongStream teaser

Abstract

Long-sequence streaming 3D reconstruction remains difficult because autoregressive visual geometry models usually anchor all poses to the first frame, which makes long-horizon prediction increasingly unstable and prone to scale drift. LongStream addresses this with a gauge-decoupled streaming formulation that predicts keyframe-relative poses, disentangles metric scale learning from geometry prediction, and periodically refreshes streaming caches to suppress long-term degradation. The resulting system supports stable metric-scale pose, depth, and point-cloud reconstruction across hundreds to thousands of frames, while remaining practical for release with full inference, evaluation, plotting, and interactive demo tooling.

ToDoList

  • Weights release
  • Model inference script
  • Minimal CLI
  • Evaluation script
  • Plotting utilities
  • Interactive demo
  • Data processing scripts release
  • Training scripts and training code release. Waiting for company approval.

Installation

Create a clean conda environment and install the runtime dependencies:

conda create -n longstream python=3.10 -y
conda activate longstream
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
conda install -c conda-forge ffmpeg -y

Notes:

  • ffmpeg is required for RGB and depth video export.
  • Please follow the official PyTorch installation guide to install Torch and CUDA versions compatible with your hardware and drivers.

Dataset Format and Conventions

LongStream expects scene data in the generalizable layout. KITTI, VKITTI 2, and Waymo should be downloaded from their official sources and converted or reorganized into this layout before running inference. Private datasets can be used directly as long as they follow the same structure and camera convention.

Official dataset links:

Generalizable Multi-Camera Layout

<meta_root>/
data_roots.txt
<scene_name>/
images/
<camera_id>/
000000.png|jpg
000001.png|jpg
...
cameras/
<camera_id>/
intri.yml
extri.yml
depths/
<camera_id>/
000000.exr
000001.exr
...

Ground-truth camera files and depth files are optional for inference. Missing ground truth only disables the corresponding evaluation metrics.

Pose and Camera Convention

LongStream uses the following convention consistently for both input annotations and saved outputs:

  • extri.yml and poses/abs_pose.txt store w2c extrinsics.
  • intri.yml and poses/intri.txt follow the OpenCV pinhole camera convention.
  • Camera axes follow OpenCV: x right, y down, z forward.
  • Depth maps are metric depths along the positive camera z axis.
  • The world frame can be arbitrary, but every camera and depth map inside one scene must share the same world frame.

Dataset Conversion

KITTI odometry:

  • Download the official KITTI odometry sequences.
  • Point --src to the KITTI root containing sequences/<seq>/image_2 and optionally image_3.
  • Convert into the generalizable meta-root:
python scripts/kitti_to_generalizable.py \
--src /path/to/kitti_odometry_root \
--out /path/to/meta_root

Waymo:

  • Export each scene into the generalizable scene layout first, with images/, cameras/, and optional depths/.
  • scripts/waymo_to_generalizable.py is a reorganization helper for an already-exported Waymo meta-root. It does not parse raw Waymo TFRecords directly.
python scripts/waymo_to_generalizable.py \
--src /path/to/waymo_meta_root \
--out /path/to/meta_root

VKITTI 2:

  • Download from the official Virtual KITTI 2 page above.
  • Export images, intrinsics, extrinsics, and optional depths into the same generalizable layout shown above.

Private datasets:

  • Export images, intrinsics, extrinsics, and optional depth directly into the generalizable layout.
  • No extra conversion script is required once the layout and camera convention match the specification above.

Checkpoints

Download the released checkpoint from the Hugging Face model repo.

You can use it in either of these ways:

  • Place 50_longstream.pt at checkpoints/50_longstream.pt.
  • Or run directly from Hugging Face:
python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--output-root outputs/seq00

Quick Start and Inference Modes

Entrypoints:

  • run.py: inference followed by evaluation
  • infer.py: inference only
  • eval.py: evaluation only on existing outputs

Inference modes:

  • batch_refresh: cuts a long sequence into overlapping keyframe spans and stitches the outputs after duplicate removal.
  • streaming_refresh: runs frame by frame with streaming caches and refreshes the caches at the same segment boundaries.

Full Run With a Local Checkpoint

python run.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--mode batch_refresh \
--streaming-mode causal \
--keyframe-stride 8 \
--refresh 3 \
--output-root outputs/

Full Run Directly From Hugging Face

python run.py \
--img-path /path/to/meta_root \
--seq-list "seq list" \
--hf-repo NicolasCC/LongStream \
--hf-file 50_longstream.pt \
--mode streaming_refresh \
--streaming-mode causal \
--window-size 48 \
--output-root outputs/

Inference Only

python infer.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--checkpoint checkpoints/50_longstream.pt \
--output-root outputs/seq00

Evaluation Only

python eval.py \
--img-path /path/to/meta_root \
--seq-list 00 \
--output-root outputs/seq00

Minimal Python API

importyamlfromlongstream.core.inferimportrun_inference_cfgwithopen("configs/longstream_infer.yaml", "r") asf:
cfg=yaml.safe_load(f)
# Point the loader to the prepared generalizable meta-root.cfg["data"]["img_path"] ="/path/to/meta_root"# Run only a subset of scenes. Remove this line to auto-discover scenes.cfg["data"]["seq_list"] = ["00"]
# Use a local checkpoint.cfg["model"]["checkpoint"] ="checkpoints/50_longstream.pt"# Or, alternatively, load the checkpoint from Hugging Face.# cfg["model"]["checkpoint"] = None# cfg["model"]["hf"] = {# "repo_id": "NicolasCC/LongStream",# "filename": "50_longstream.pt",# }# Select the output directory for saved predictions.cfg["output"]["root"] ="outputs/seq00_py"# This writes poses, depths, point clouds, RGB/depth videos, and visualizations to disk.run_inference_cfg(cfg)

run_inference_cfg saves predictions under cfg["output"]["root"]. The saved artifacts include per-frame poses and intrinsics, raw depth maps, colorized depth visualizations, RGB frame exports, and merged point clouds for both released branches.

Important Runtime Args

ArgMeaning
--modebatch_refresh or streaming_refresh
--streaming-modecausal or window attention
--window-sizeattention window for window mode
--keyframe-stridekeyframe interval
--refreshkeyframes per refresh span
--cameraselect one camera in a multi-camera scene
--mask-skyenable sky masking; requires onnxruntime and may auto-download skyseg.onnx
--no-mask-skydisable sky masking, useful for offline runs or deterministic packaging
--skip-evalskip evaluation in run.py

Demo

Install the demo requirements first:

To let the demos fetch the released checkpoint automatically from Hugging Face:

export LONGSTREAM_HF_REPO=NicolasCC/LongStream
export LONGSTREAM_HF_FILE=50_longstream.pt

Stable demo:

python demo_gradio.py

Outputs

For each sequence, LongStream writes a dedicated directory under output_root/<sequence>/:

  • poses/abs_pose.txt: per-frame absolute w2c extrinsics.
  • poses/rel_pose.txt: per-frame relative pose-head outputs when the relative pose head is enabled.
  • poses/intri.txt: per-frame intrinsics.
  • images/rgb/ and images/rgb.mp4: RGB frame exports and video.
  • depth/dpt/: raw predicted depth maps as .npy.
  • depth/dpt_plasma/ and depth/dpt_plasma.mp4: colorized depth previews.
  • points/point_head*: point clouds from the point-head branch.
  • points/dpt_unproj*: point clouds obtained by unprojecting predicted depth with predicted poses.

When evaluation is run, metrics are saved to:

  • output_root/metrics/<sequence>.json
  • output_root/summary.json
  • output_root/plots/<sequence>_traj_3d.png

Citation

@misc{cheng2026longstreamlongsequencestreamingautoregressive,
title={LongStream: Long-Sequence Streaming Autoregressive Visual Geometry}, author={Chong Cheng and Xianda Chen and Tao Xie and Wei Yin and Weiqiang Ren and Qian Zhang and Xiaoyang Guo and Hao Wang},
year={2026},
eprint={2602.13172},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.13172}, }

Acknowledgements

We thank the authors and open-source projects of VGGT, Stream3R, StreamVGGT, DUSt3R, and CroCo for releasing code that informed this inference release.

About

No description, website, or topics provided.

Resources

Stars

94 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages