Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting 🚀🤖

SlidePaperHugging Face

Overview

VOTE is a framework for building fast and accurate Vision-Language-Action (VLA) models for robotic manipulation. It introduces Trajectory Ensemble Voting, a method that optimizes action prediction by ensembling multiple trajectory candidates decoded from a vision-language backbone. VOTE achieves state-of-the-art performance on both simulated (LIBERO, SimplerEnv) and real-world benchmarks while being 3× faster than prior VLA methods. Its modular design allows easy migration to any VLM backbone with just 2 lines of code—no complex action tokenizers required.

Table of Contents

News

  • 2025/09/22: ✨ Released VOTE LLAMA3.2-1B-VLA model 👉 script — inference with only 4.34 GB VRAM usage.
  • 2025/07/10: 🎉 Released VOTE 1.0. ➡️ No need for complex tokenizers — migrate to a new VLM with just 2 lines of code ⚡️

Installation

conda create -n vote python=3.10 -y
conda activate vote
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .
pip install packaging ninja
ninja --version;echo$?# Verify Ninja --> should return exit code "0"
pip install flash-attn==2.6.1 --no-build-isolation

Quick Start

cd experiments/speed/
python effvla.py

Speed Benchmarks

We provide speed measurement scripts under experiments/speed/:

# OpenVLA-7B (2-token, 16-action trajectory ensemble)
python experiments/speed/effvla.py
# LLAMA3.2-1B-VLA (lightweight, ~4.34 GB VRAM)
python experiments/speed/llama3-1B.py
# Other baselines
python experiments/speed/openvla.py
python experiments/speed/cogact.py
python experiments/speed/pi0.py
python experiments/speed/spatialvla.py

Troubleshooting

No module named prismatic / No module named experiments

This usually means the package was not installed correctly. Verify with:

pip list | grep effvla

If effvla is not listed, re-run pip install -e . from the repo root.

If you run into any other issues, please open a GitHub issue.

Installation on Jetson AGX Orin

Click to expand
python -m venv orin
source orin/bin/activate
# Install transformers and other dependencies
pip3 install packaging ninja transformers==4.51.0 tokenizers==0.21.4 timm==0.9.10 diffusers==0.32.2
# Install TensorFlow 2.15.0
pip3 install tensorflow==2.15.0
# Install TensorFlow addons from source
git clone https://github.com/tensorflow/addons
cd addons
pip3 install -e .cd ..
# Install VOTE
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip3 install -e .cd ..
# Note: This step installs torch/torchvision versions incompatible with Jetson.# We will override them with precompiled wheels below.# Install torch & torchvision using NVIDIA's precompiled wheels for Jetson# torch: https://nvidia.box.com/shared/static/mp164asf3sceb570wvjsrezk1p4ftj8t.whl# torchvision: https://nvidia.box.com/shared/static/xpr06qe6ql3l6rj22cu3c45tz1wzi36p.whl
pip3 install torch*.whl torchvision*.whl

Note: You may see a dependency conflict warning like
effvla 0.0.1 requires torchvision==0.18.1, but you have torchvision 0.18.0a0+6043bc2.
This is expected and can be safely ignored.

Architecture

Multi-Token Trajectory Ensemble

VOTE replaces standard single-token action decoding with a multi-token trajectory ensemble approach. The VLM backbone generates multiple <ACT> tokens, each decoded into an action chunk by a lightweight action head. The final trajectory is assembled by ensembling predictions across tokens.

Key parameters:

ParameterDescriptionExample
num_actions_chunkTotal actions in the output sequence8 or 16
num_actions_per_tokenActions predicted per <ACT> token8
modePrediction mode ("mul" for multi-token ensemble)"mul"

For example, num_actions_chunk=16 with num_actions_per_token=8 produces 2 tokens, each predicting 8 actions (16 total).

Action Heads

Three action head architectures are available via --action_head_name:

NameClassDescription
mlpL1RegressionActionHeadmulmlpkStandard MLPResNet-based head
felL1RegressionActionHeadFunnelFunnel architecture with progressive dimension reduction (more parameter-efficient)

Additional parameters:

  • --num_blocks: Number of MLPResNet blocks (typically 2 for Fractal, 4 for LIBERO)
  • --hidden_dim: Hidden dimension size (default: 4096 for OpenVLA-7B, 2048 for LLAMA3.2-1B)

Supported Backbones

BackboneBase Model Pathmodel_typeParams
OpenVLA-7B (LLaMA 2)openvla/openvla-7bllama27B
LLAMA3.2-1B-VLAjuyil/llama3.2-1B-VLMllama3.22.3B

Training

Training Environment

Training runs on NVIDIA H100 NVL GPUs (94 GB VRAM each) with 756 GB RAM. We use a shuffle buffer of 256K samples.

Data Preparation

BridgeDataV2 and Fractal are part of the Open X-Embodiment dataset. Follow rlds_dataset_mod for data preparation.

Running Training

Fractal (single GPU):

bash train.sh

The default configuration uses:

torchrun --standalone --nnodes 1 --nproc-per-node 1 vla-scripts/train.py \
--vla_path openvla/openvla-7b \
--data_root_dir /data/ \
--dataset_name fractal20220817_data \
--run_root_dir /data/wandbrun \
--use_l1_regression True \
--batch_size 4 \
--learning_rate 1e-4 \
--max_steps 200005 \
--save_freq 5000 \
--image_aug True \
--lora_rank 32 \
--num_actions_chunk 8 \
--num_actions_per_token 8 \
--num_blocks 2 \
--mode "mul" \
--action_head_name "funnel"

For LIBERO training, refer to LIBERO.md.

Training Parameters

ParameterDescriptionDefault
--vla_pathBase VLM checkpoint (HF Hub or local)openvla/openvla-7b
--use_l1_regressionUse L1 regression action headTrue
--use_diffusionUse diffusion-based action head (DDIM)False
--use_filmUse FiLM language-vision conditioningFalse
--use_proprioInclude proprioceptive state in inputFalse
--num_images_in_inputNumber of camera images (1 = 3rd person only)1
--lora_rankLoRA rank for fine-tuning32
--image_augEnable random crop image augmentationTrue
--num_actions_chunkTotal action sequence length
--num_actions_per_tokenActions decoded per token
--num_blocksMLPResNet depth in action head
--modePrediction mode ("mul" for ensemble)"mul"
--action_head_nameAction head type ("mlp", "funnel", "fel")"funnel"

Evaluation

LIBERO

Follow LIBERO.md for LIBERO setup, training, and evaluation.

Quick multi-GPU evaluation:

# Using the shell launcher (recommended)
CKPT_DIR=/path/to/ckpts TASK_SUITE=libero_goal bash run_libero_goal_eval.sh
# Using the Python script
python experiments/robot/libero/batch_eval.py \
--dir /path/to/ckpts \
--task_suite libero_goal \
--devices 0 1 2 3 4 5 6 7

SimplerEnv

Important: Install SimplerEnv before installing effvla, because installing TensorFlow 2.15 may break the CUDA environment for PyTorch.

SimplerEnv installation steps
conda create -n simpler_env python=3.10
conda activate simpler_env
git clone https://github.com/LukeLIN-web/simplerenv.git --recurse-submodules
pip install numpy==1.24.4 # numpy >= 1.26 causes issues in SimplerEnvcd simplerenv/ManiSkill2_real2sim
pip install -e .cd ..
pip install -e .cd ..
git clone https://github.com/LukeLIN-web/vote.git
cd vote
pip install -e .cd ..
sudo apt install ffmpeg
cd simplerenv
pip install tensorflow==2.15.0
pip install "tensorflow[and-cuda]==2.15.1"# TensorFlow GPU support# If you encounter: libtorch_cuda.so: undefined symbol: ncclCommRegister# Re-install torch and torchvision:
pip install torch==2.3.1 torchvision==0.18.1
pip install mediapy pandas gymnasium==0.28.1

Results

SimplerEnv (WidowX, Visual Matching)

MethodPut SpoonPut CarrotStack BlockPut EggplantAvg.Latency (ms) ↓Speedup ↑
RT-1-X0.04.20.00.01.1
Octo47.29.74.256.930.0
OpenVLA0.00.00.04.11.02401.00
RoboVLM29.225.012.558.331.3
OpenPI029.10.016.662.527.14700.50
SpatialVLA16.725.029.2100.042.74000.60
CogACT71.750.815.067.551.32201.09
VOTE (Ours)58.329.250.095.858.3783.1

LLAMA3.2-1B-VLA

ModelParams (B)libero_spatiallibero_objectlibero_goallibero_10Average SRVRAM (GB)
LLAMA3.2-1B-VLA2.398.4%96.0%95.0%82.4%92.95%4.34

Edge deployment latency:

  • Jetson AGX Orin: 108 ms (chunk = 8, ≈ 73 Hz)
  • Jetson Nano: 387 ms

Citation

If you find this work useful, please cite our paper:

@misc{lin2025votevisionlanguageactionoptimizationtrajectory,
title={VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting}, author={Juyi Lin and Amir Taherin and Arash Akbari and Arman Akbari and Lei Lu and Guangyu Chen and Taskin Padir and Xiaomeng Yang and Weiwei Chen and Yiqian Li and Xue Lin and David Kaeli and Pu Zhao and Yanzhi Wang},
year={2025},
eprint={2507.05116},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05116}, }

License

This project is released under the MIT License.

About

Vision-Language-Action Optimization with Trajectory Ensemble Voting (ICANN2026)

Resources

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages