Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

Accepted at CVPR 2026

arXivHugging Face

If you find this repository helpful, a star would be greatly appreciated.

Abstract

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised.

To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion.

Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance.

Framework

CCF framework

Main Results

CCF results

Reproduction

This README is intended to be runnable end-to-end for reproduction. A coding agent can use it to set up the environment, download the released assets, and run the split evaluation; we tested this workflow with Codex using GPT-5.5.

Environment

The release uses a Torch 2 based MMDetection/MMDetection3D stack provided through submodules.

Clone with submodules, or initialize them after cloning:

git submodule update --init --recursive

The required third-party repositories are tracked as submodules:

thirdparty/mmcv_torch2
thirdparty/mmdetection_ccf
thirdparty/mmdetection3d_ccf
thirdparty/nuscenes-devkit_ccf

A setup script is provided as a reference for the installation sequence tested on NVIDIA RTX 5090:

conda create -n ccf python=3.10 -y
conda activate ccf
bash setup.sh

If your CUDA or driver stack differs, adjust the PyTorch and spconv wheels accordingly.

Data

Place the official nuScenes data under data/nuscenes/ with the standard layout:

data/nuscenes/
├── maps/
├── samples/
├── sweeps/
└── v1.0-trainval/

If nuScenes already exists elsewhere, using a symlink is sufficient.

CCF also needs generated info files and source/target split pkl files. They can be downloaded from the Hugging Face assets repo:

pip install -U huggingface_hub
hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "data/nuscenes/nuscenes_infos*.pkl""data/nuscenes/splits/*"

The CCF configs expect these source split files:

data/nuscenes/splits/nuscenes_infos_train_singapore_norain_day_source.pkl
data/nuscenes/splits/nuscenes_infos_val_singapore_norain_day_source.pkl

The split evaluation configs expect:

data/nuscenes/splits/nuscenes_infos_val_night_target.pkl
data/nuscenes/splits/nuscenes_infos_val_rain_target.pkl
data/nuscenes/splits/nuscenes_infos_val_boston_target.pkl

The released assets already include the required source and target split files. To create splits locally instead, first regenerate the full nuScenes info files:

python tools/create_data_nusc.py \
--root-path data/nuscenes \
--version v1.0 \
--extra-tag nuscenes \
--max-sweeps 10

Then export the desired source/target split pkl files from the full nuscenes_infos*.pkl files with the scripts in tools/nuscenes_data_split/; see tools/nuscenes_data_split/README.md for the exact commands.

Checkpoints

Download the released checkpoints and initialization weights:

hf download Curiosity-Wu/CCF \
--repo-type model \
--local-dir . \
--include "checkpoints/*.pth"
mkdir -p checkpoints/eval
ln -sf ../ccf_source.pth checkpoints/eval/ccf_source.pth

The source-domain training configs expect two initialization checkpoints:

checkpoints/isfusion_source.pth
checkpoints/faster_rcnn_swint_fpn_source.pth

Both source initialization checkpoints are trained on the nuScenes source split. faster_rcnn_swint_fpn_source.pth is trained on 2D boxes obtained by projecting nuScenes 3D boxes from the source split onto images.

The oracle training config expects the corresponding full-train initialization checkpoints:

checkpoints/isfusion_oracle.pth
checkpoints/faster_rcnn_swint_fpn_oracle.pth

Split evaluation expects:

checkpoints/eval/ccf_source.pth

Training

Train the CCF model:

bash tools/dist_train.sh projects/configs/ccf/ccf_source.py 4

Train the source baseline:

bash tools/dist_train.sh projects/configs/ccf/baseline_source.py 4

Train the oracle model:

bash tools/dist_train.sh projects/configs/ccf/ccf_oracle.py 4

The oracle config trains on the full nuScenes train split and validates on the full nuScenes val split, so it is an upper-bound setting rather than the source-domain reproduction setup. It uses depth_loss_type="silog" for the image depth branch to improve training stability.

Evaluation

Evaluate the CCF checkpoint on the configured target splits. The reproduction run above used one GPU. If multiple GPUs are available, prefer setting GPUS to the number of usable GPUs for faster evaluation:

GPUS=4 bash tools/eval_splits.sh

For a single-GPU run, use:

GPUS=1 bash tools/eval_splits.sh

The individual evaluation configs live under projects/configs/ccf/eval/:

projects/configs/ccf/eval/ccf_source-night.py
projects/configs/ccf/eval/ccf_source-rain.py
projects/configs/ccf/eval/ccf_source-boston.py

Notes on ISFusion Modifications

We modified the ISFusion detector used in this repository to obtain empirically better results in the CCF reproduction setting. The main changes from the original ISFusion implementation are:

  1. The detection head decouples classification and regression. The decoder now produces separate classification and box features, and the prediction head uses the classification feature for heatmap prediction and the box feature for box regression.
  2. The head adds a center refinement step before the final decoupled prediction. With the released one-layer decoder setting, the original ISFusion head predicts the final center as an offset from the initial proposal center. This version first refines the proposal center with a dedicated center decoder/head, then uses the refined center for the final classification and box prediction.

Notes on Training Stability

The released training configs include two stability-oriented implementation choices introduced during CCF development: SigmaReparam, following apple/ml-sigma-reparam, and CAdamW, following hazdzz/c_adam. We adopted them to mitigate occasional loss spikes during training.

Citation

If you find CCF useful for your research, please cite:

@InProceedings{Wu_2026_CVPR,
author = {Yuchen Wu and Kun Wang and Yining Pan and Na Zhao},
title = {CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {18745-18754}
}

Acknowledgement

This codebase builds on MMDetection3D, MV2DFusion and ISFusion.

About

No description, website, or topics provided.

Resources

Stars

19 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages