Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

ACL 2025 | PruneVid

The official repository for paper "PruneVid: Visual Token Pruning for Efficient Video Large Language Models".

Xiaohu Huang, Hao Zhou, Kai Han

WebpagePaper

Introduction

Framework

We present PruneVid, a training-free visual token pruning method that enhances efficiency in multi-modal video understanding. By merging spatial-temporal tokens to reduce video redundancy and leveraging attention mechanisms within LLMs to retain only the visual tokens relevant to questions, PruneVid ensures high performance while reducing computational overhead.

Todo:

  • Code release of PruneVid with PLLaVA.
  • Code release of PruneVid with LLaVA-OneVision.
  • Code release of PruneVid with ST-LLM.

License

PruneVid is released under the CC BY-NC-SA 4.0 license.

Performance

We conduct experiments on three video LLMs (PLLaVA, ST-LLM, and LLaVA-OneVision) under for benchmarks: MVBench, VideoMME, Egoschema, and VideoChatgpt-Bench (VCG-Bench).

MethodRetained RatioFLOPs (×)MVBenchVideoMMEEgoSchema Subset / FullsetTUCUCODOCIAvg
PLLaVA100.0%1.00×46.644.447.8 / 42.62.333.622.932.863.212.99
PLLaVA w/ FastV30.0%0.33×46.143.646.2 / 41.02.383.492.892.763.142.93
PLLaVA w/ Prumerge55.7%0.53×45.643.845.2 / 40.42.343.522.902.763.152.93
PLLaVA w/ Look-M20.0%1.00×46.644.347.0 / 42.32.283.412.752.653.002.82
PLLaVA w/ Ours16.2%0.23×47.645.349.0 / 42.62.443.512.992.783.202.98
ST-LLM100.0%1.00×54.942.056.2 / 45.62.463.462.662.633.082.86
ST-LLM w/ FastV30.0%0.37×42.934.548.0 / 38.52.012.231.551.941.691.88
ST-LLM w/ Look-M20.0%1.00×54.040.654.0 / 44.52.353.412.602.513.012.78
ST-LLM w/ Ours15.1%0.26×54.341.454.6 / 44.72.403.432.632.603.042.82
LLaVA-OneVision100.0%1.00×58.058.262.0 / 60.02.753.703.392.973.503.26
LLaVA-OneVision w/ FastV30.0%0.30×57.257.662.6 / 60.02.653.613.282.853.393.16
LLaVA-OneVision w/ Prumerge55.2%0.49×52.956.762.2 / 60.02.723.643.322.943.443.21
LLaVA-OneVision w/ Look-M20.0%1.00×57.058.062.0 / 59.82.713.703.292.893.443.21
LLaVA-OneVision w/ Ours17.0%0.20×57.558.662.6 / 59.52.733.723.282.943.513.24

Data Preparation

All four used benchmarks can be downloaded from huggingface website: MVBench, VideoMME, Egoschema, and VideoChatGPT-Bench.

After downloading the datasets, please put them into the DATAS folder and sort out the source videos and annotations in the following formats:

DATAS/
├── ego_schema/
│ ├── json/
│ └── videos/
├── MVBench/
│ ├── json/
│ └── video/
├── VCGBench/
│ ├── Videos/
│ ├── Zero_Shot_QA/
└── Video-MME/
├── data/
└── json/

Pretrained Model

The pretrained model can be found in their respective repositories: PLLaVA, ST-LLM, and LLaVA-OneVision.

After downloading the models please put them into the MODELS folder:

MODELS/
├── pllava-7b/

Environment Install

We follow the environment installation guideline of PLLaVA.

  1. Above all, the following environment set up is for python 3.10. If you choose to use conda for environment setup, we recommend creating the virtual environment with:
conda create -n pllava python=3.10
  1. Firstly, install pytorch from the official website. The code runs on torch 2.2.1, cu118 or cu122. Select the version that suits your drive version.
torch 2.2.1+cu118
torchaudio 2.2.1+cu118
torchvision 0.17.1+cu118

If your driver version is higher than cu121, you could probably try installing with the following scripts:

pip install -r requirements.txt

Otherwise, you would need to install a torch for your server first, then install the other packages:

pip install -r requirements.torch.txt # decide your own requirements, (this is for cu11), or install torch directly following the official website.
pip install -r requirements.no_torch.txt # install the following

Evaluation

As PruneVid is a training-free method, we can directly apply it on the pre-trained models.

The provided scripts for evaluating model performance is given in scripts/eval.sh. Below is the script for evaluating the performance on MVBench, where you can edit the hyper-parameters whatever you want. The default setting is used in our paper.

lora_alpha=14
selected_layers=(10)
alphas=(0.4)
taus=(0.8)
temporal_segment_ratios=(0.25)
cluster_ratios=(0.5)
for alpha in "${alphas[@]}"; do
for selected_layer in "${selected_layers[@]}"; do
for tau in "${taus[@]}"; do
for temporal_segment_ratio in "${temporal_segment_ratios[@]}"; do
for cluster_ratio in "${cluster_ratios[@]}"; do
# 执行命令
SAVE_DIR=test_results/pllava-7b-lora${lora_alpha}-threshold${tau}-layer${selected_layer}-alpha${alpha}-temporal-segment-ratio-${temporal_segment_ratio}-cluster-ratio-${cluster_ratio}
mkdir -p "${SAVE_DIR}"
conv_mode=eval_mvbench
python -m tasks.eval.mvbench.pllava_eval_mvbench \
--pretrained_model_name_or_path ${model_dir} \
--save_path ${SAVE_DIR}/mvbench \
--num_frames ${num_frames} \
--use_lora \
--lora_alpha ${lora_alpha} \
--top_p 1.0 \
--temperature 1.0 \
--weight_dir ${weight_dir} \
--pooling_shape 16-12-12 \
--conv_mode ${conv_mode} \
--selected_layer ${selected_layer} \
--alpha ${alpha} \
--tau ${tau} \
--temporal_segment_ratio ${temporal_segment_ratio} \
--cluster_ratio ${cluster_ratio}
done
done
done
done
done

As for Egoschema, which needs an external service to evaluate the model performance, we run the evaluate_egoschema_result.py for evaluation. Before executing the file, you should change the root_dir variable to your folder.

python evaluate_egoschema_result.py

Acknowledgement

This repository is built upon PLLaVA, ST-LLM, and LLaVA-OneVision. Thanks for those well-organized codebases.

Citation

@inproceedings{
huang2024prunevid,
title={PruneVid: Visual Token Pruning for Efficient Video Large Language Models},
author={Xiaohu Huang and Hao Zhou and Kai Han},
booktitle={arXiv},
year={2024}
}

About

[ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Resources

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages