Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - HiThink-Research/CCPO: Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents · GitHub
Skip to content

Latest commit

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Project PagearXivLicense: MITPython 3.12+Pytorch 2.7+Hugging Face

Yurun Song*, Jiong Yin*, Rongjunchen Zhang, Ian Harris

🔥 News

  • [2026-08-07] Our paper has been accepted to ACM MM 2026 Main Track, see you in Brazil!
  • [2026-01-12] 🚀 Code and pre-trained models are released!
  • [2026-01-14] 📄 Our paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents" is now available on arXiv.

🚀 Introduction

The official implementation of the paper "Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents".

Abstract:Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

📈 Method Overview

Method OverviewOverview of the CCPO framework.

✨ Key Features

  • Efficient Compression (CASC): Aggregates spatial coordinates to achieve up to 60% token reduction without losing critical context.
  • Distance-Based Advantage: Provides fine-grained learning signals based on spatial distance, significantly boosting grounding accuracy.
  • Training Acceleration: Delivers 3.5x–4.8x speedup and 16% lower TFLOPS compared to standard RL baselines.
  • SOTA Performance: Top-tier results across 4 major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW.
  • Coupled Optimization: A unified framework that co-optimizes visual focusing and policy decision-making.

🛠️ Installation

Requirements

  • Linux
  • Python 3.12+
  • PyTorch 2.7+
  • CUDA 12.8+
  • Please refer to requirements.txt for other dependencies.

Setup

# Clone the repository
git clone https://github.com/HiThink-Research/CCPO.git
cd CCPO
# Create a conda environment
conda create -n ccpo python=3.12
conda activate ccpo
# Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

📂 Data Preparation

We evaluate CCPO on four major benchmarks: Android Control, GUI Odyssey, Mind2Web, and AITW, please organize the data as follows:

data/
├── android_control/
├── gui_odyssey/
├── mind2web/
└── aitw/

🏃 Usage

1. SFT Training

We first perform Supervised Fine-Tuning (SFT) on Qwen2.5-VL as the warm-up stage.

2. CCPO Training

Then we train the CCPO model with the following command:

cd CCPO
bash scripts/train_CCPO_aitw_7B.sh

2. Evaluation

To evaluate the pre-trained model:

cd ../evaluation
python evaluation_aitw.py \
--save_path path/to/save/results \
--model_path path/to/model \
--his_num 4

📊 Model Zoo

We provide pre-trained models (3B and 7B) for reproduction.

DatasetCCPO-3BCCPO-7B
AITWDownloadDownload

📝 Citation

If you find our work useful for your research, please consider citing:

@misc{song2026compressfocusefficientcoordinate,
title={Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents}, author={Yurun Song and Jiong Yin and Rongjunchen Zhang and Ian G. Harris},
year={2026},
eprint={2601.11631},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.11631}, }

🙏 Acknowledgement

This project is built upon UI-S1, SimpAgent, and verl-agent. We thank the authors for their great code.

About

Compress2Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages