Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - THUSI-Lab/Hstar: [CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild · GitHub
Skip to content

Repository files navigation

Thinking in 360°: Humanoid Visual Search in the Wild

*Equal contributionCorresponding author
1New York University2NVIDIA3TU Darmstadt4UC Berkeley5Stanford University
arXivWebsiteHF Model: HVS-3BHF Dataset: hvs-train-datasetsHF Dataset: hstar_benchmark

teaser

News

  • [2025/2/21] Our paper is accepted by CVPR 2026!
  • [2025/2/15] Updated benchmarking dataset with high-resolution .png files at here.
  • [2025/11/26] Our paper is available on arXiv.
  • [2025/11/26] We release our finetuend HVS-3B model on HuggingFace.
  • [2025/11/26] We release our training datasets on HuggingFace.
  • [2025/11/26] We release our benchmarking dataset on HuggingFace.

Getting Started

Installation

Set up the VAGEN environment for training.

conda create -n vagen python=3.10
conda activate vagen
git clone --recursive https://github.com/humanoid-vstar/hstar.git
cd hstar
cd verl && pip install -e .cd ..
bash scripts/install.sh

For benchmarking, we need a different envrionment for later transformers and vllm version.

conda create -n hstar python=3.10
conda activate hstar
cd vagen/inference && pip install -r requirements.txt # This env is build for CUDA 12 and torch 2.7.1# You need to adjust the environment to adapt your machine.# for new models, use a vllm env to deploycd ../..
cd verl && pip install -e . --no-deps
cd .. && pip install -e .

In addition, if you want to train the model from scratch, you need to install LLaMA-Factory for SFT training.

Training

SFT

Use LLaMA-Factory and our SFT dataset hos_sft and hps_sft to train or directly download our fine-tuned model HVS-3B-sft-only.

RL

  • Download our RL dataset hvs_rl (use mixed_rl.zip if you want to trained on the mixed dataset)
  • Change your downloaded dataset path in the training config.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkdata_path: /path/to/your/datasetuse_state_reward: falsetraj_success_reward: 0.5traj_fail_penalty: 0format_reward: 0.5resolution: 720train_size: 3200test_size: 32
  • Change your model path in scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh or the tmux-free script and modify other hyperparameters.
    # ...
    actor_rollout_ref.model.path=/path/to/your/model \\# ...
    critic.model.path=/path/to/your/model \\# ...
  • Then run the experiment by:
    # With tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run_tmux.sh
    # Without tmux
    bash scripts/examples/masked_grpo/hstar/free_think/run.sh

Benchmarking

  • Download our hstar_bench dataset.
  • Change your downloaded dataset path (2 task splits) in the scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml and scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml.
    env1:
    env_name: hstar env_config:
    render_mode: visionprompt_format: free_thinkuse_state_reward: falsedata_path: /path/to/your/dataset/splitresolution: 1080
  • Create test dataset seeds.
    # Create one full dataset
    python vagen/env/create_dataset.py
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench/train.parquet" \
    --test_path "data/hos_bench/test.parquet"
    python vagen/env/create_dataset.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench/train.parquet" \
    --test_path "data/hps_bench/test.parquet"# Or dataset clips for better efficiency
    python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hos_val_config.yaml" \
    --train_path "data/hos_bench_clip/train.parquet" \
    --test_path "data/hos_bench_clip/test.parquet" \
    --num_clip 10 python vagen/env/create_dataset_clip.py \
    --yaml_path "scripts/examples/masked_grpo/hstar/free_think/hps_val_config.yaml" \
    --train_path "data/hps_bench_clip/train.parquet" \
    --test_path "data/hps_bench_clip/test.parquet" \
    --num_clip 10
  • Modify inference settings in vagen/inference/inf_cfg.yaml and model settings in vagen/inference/model_cfg.yaml
  • Deploy your model using vllm OpenAI API Server on localhost:8000, see example vagen/inference/deploy.sh
  • Run the experiment
    cd vagen/inference
    python -m vagen.server.server server.port=5000 &> ./inf_server.log &
    python run_inference.py \
    --inference_config_path inf_cfg.yaml \
    --model_config_path model_cfg.yaml \
    --val_files_path /path/to/your/generated/seeds/path \
    --wandb_path_name hstar_bench &
    [--output_dir /path/to/output/dir] # default ./temp_result
    [--save_all_results False] # save all the ouputs when set to True 
  • View result
    python show_result.py [--result_dir /path/to/output/dir] # default ./temp_result

References and Acknowledgement

  • LLaMA-Factory: Easy and Efficient LLM Fine-Tuning

  • VAGEN: VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents

  • verl: Volcano Engine Reinforcement Learning for LLM

Citation

@misc{yu2025thinking360deghumanoidvisual,
title={Thinking in 360°: Humanoid Visual Search in the Wild}, author={Heyang Yu and Yinan Han and Xiangyu Zhang and Baiqiao Yin and Bowen Chang and Xiangyu Han and Xinhao Liu and Jing Zhang and Marco Pavone and Chen Feng and Saining Xie and Yiming Li},
year={2025},
eprint={2511.20351},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.20351}, }

About

[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild

Topics

Resources

Stars

146 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages