Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SSync | ECCV 2026

Selective Synergistic Learning for Video Object-Centric Learning

WonJun Moon1 · Jae-Pil Heo2
1KAIST 2Sungkyunkwan University

PaperarXivLicenseHugging Face ModelProject Page

Official PyTorch implementation of SSync (Selective Synergistic Learning), a selective mutual-distillation framework for video object-centric learning.


📖 Abstract

Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder–decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit different properties, a recent dense alignment strategy attempted to reconcile this discrepancy by enforcing agreement across all spatio-temporal patches via contrastive learning. However, this indiscriminate alignment inadvertently propagates the inherent weaknesses of each module, such as noisy encoder predictions and blurred decoder boundaries. Moreover, computing dense similarities across all pairs incurs a computational cost quadratic in the total number of spatio-temporal patches, severely limiting scalability.

Motivated by this, we propose Selective Synergistic Learning (SSync). Instead of exhaustive patch-to-patch alignment, SSync prevents error propagation by selectively distilling only the most reliable cues: leveraging the encoder strictly for boundary refinement and the decoder for interior denoising. This is realized via pseudo-labeling with linear complexity, eliminating the need for quadratic spatial comparisons. Also, to prevent the reinforcement of architectural biases like slot redundancy, we introduce a transitive pseudo-label merging that consolidates overlapping slots based on spatio-temporal activation consistency. Extensive studies demonstrate that SSync improves decomposition quality and serves as a versatile, plug-and-play module while also exhibiting exceptional robustness to slot configurations.

SSync selective supervision (Figure 2)


⚙️ Installation

pip install poetry
poetry lock
poetry install
poetry run pip install matplotlib coco notebook
poetry run pip install tensorboard tensorboardX

🗂️ Datasets

To download the datasets used in this work, see the instructions in data/README.md. For more details, we refer to the SlotContrast repository.

The datasets should be placed under a common root directory with the following structure:

├── SSync/
└── dataset/
├── ytvis2021_resized/
├── movi_c/
└── movi_e/
DatasetDownloadSize
YouTube-VIS 2021Google Drive26.43 GB
MOVi-CGoogle Drive7.43 GB
MOVi-EGoogle Drive8.26 GB

🚀 Training

Run one of the configurations in configs/SSync, for example:

poetry run python -m SSync.train --run-eval-after-training configs/SSync/coco.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_c.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/movi_e.yaml
poetry run python -m SSync.train --run-eval-after-training configs/SSync/ytvis2021.yaml

🧪 Pretrained Checkpoints

DatasetDownload
MOVi-CGoogle Drive
MOVi-EGoogle Drive
YouTube-VIS 2021Google Drive
COCO 2017Google Drive

🖼️ Qualitative Results

Each clip blends the original input with the predicted slot map (each object in a distinct color). For interactive controls, visit the project page.

MOVi-CMOVi-EYouTube-VIS 2021

📌 Citation

If you find this work useful, please consider citing:

@inproceedings{moon2026ssync,
title = {Selective Synergistic Learning for Video Object-Centric Learning},
author = {Moon, WonJun and Heo, Jae-Pil},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}

as well as SlotCurri and SRL:

@inproceedings{moon2026reconstruction,
title = {Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning},
author = {Moon, WonJun and Seong, Hyun Seok and Heo, Jae-Pil},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
@inproceedings{seong2026synergistic,
title = {From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning},
author = {Seong, Hyun Seok and Moon, WonJun and Heo, Jae-Pil},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}

🙏 Acknowledgement

Our implementation is built upon the official repositories of VideoSAUR, SlotContrast, SRL, and SlotCurri.


📄 License

This codebase is released under the MIT License. Some parts of the codebase were adapted from other codebases; a comment was added to the code where this is the case, and those parts are governed by their respective licenses.