Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

98 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Paper2Video

English | 简体中文

Paper2Video: Automatic Video Generation from Scientific Papers
从学术论文自动生成演讲视频

Zeyu Zhu*, Kevin Qinghong Lin*, Mike Zheng Shou
Show Lab, National University of Singapore

📄 Paper | 🤗 Daily Paper | 📊 Dataset | 🌐 Project Website | 💬 X (Twitter)

  • Input: a paper ➕ an image ➕ an audio
PaperImageAudio

🔗 Paper link

Hinton's photo

🔗 Audio sample
  • Output: a presentation video
hinton.2.mp4

Check out more examples at 🌐 project page.

🔥 Update

Any contributions are welcome!

  • [2025.10.15] We update a new version without talking-head for fast generation!
  • [2025.10.11] Our work receives attention on YC Hacker News.
  • [2025.10.9] Thanks AK for sharing our work on Twitter!
  • [2025.10.9] Our work is reported by Medium.
  • [2025.10.8] Check out our demo video below!
  • [2025.10.7] We release the arxiv paper.
  • [2025.10.6] We release the code and dataset.
  • [2025.9.28] Paper2Video has been accepted to the Scaling Environments for Agents Workshop(SEA) at NeurIPS 2025.
d35df30ad813f1cf53eccab0ad525b5d.mp4

Table of Contents


🌟 Overview

Overview

This work solves two core problems for academic presentations:

  • Left: How to create a presentation video from a paper?
    PaperTalker — an agent that integrates slides, subtitling, cursor grounding, speech synthesis, and talking-head video rendering.

  • Right: How to evaluate a presentation video?
    Paper2Video — a benchmark with well-designed metrics to evaluate presentation quality.


🚀 Try PaperTalker for your Paper!

Approach

1. Requirements

Prepare the environment:

cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic

[Optional] Skip this part if you do not need a human presenter.

Download the dependent code and follow the instructions in Hallo2 to download the model weight.

git clone https://github.com/fudan-generative-vision/hallo2.git

You need to prepare the environment separately for talking-head generation to potential avoide package conflicts, please refer to Hallo2. After installing, use which python to get the python environment path.

cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt

2. Configure LLMs

Export your API credentials:

export GEMINI_API_KEY="your_gemini_key_here"export OPENAI_API_KEY="your_openai_key_here"

The best practice is to use GPT4.1 or Gemini2.5-Pro for both LLM and VLMs. We also support locally deployed open-source model(e.g., Qwen), details please referring to Paper2Poster.

3. Inference

The script pipeline.py provides an automated pipeline for generating academic presentation videos. It takes LaTeX paper sources together with reference image/audio as input, and goes through multiple sub-modules (Slides → Subtitles → Speech → Cursor → Talking Head) to produce a complete presentation video. ⚡ The minimum recommended GPU for running this pipeline is NVIDIA A6000 with 48G.

Example Usage

Run the following command to launch a fast generation (without talking-head generation):

python pipeline_light.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--gpu_list [0,1,2,3,4,5,6,7]

Run the following command to launch a full generation (with talking-head generation):

python pipeline.py \
--model_name_t gpt-4.1 \
--model_name_v gpt-4.1 \
--model_name_talking hallo2 \
--result_dir /path/to/output \
--paper_latex_root /path/to/latex_proj \
--ref_img /path/to/ref_img.png \
--ref_audio /path/to/ref_audio.wav \
--talking_head_env /path/to/hallo2_env \
--gpu_list [0,1,2,3,4,5,6,7]
ArgumentTypeDefaultDescription
--model_name_tstrgpt-4.1LLM
--model_name_vstrgpt-4.1VLM
--model_name_talkingstrhallo2Talking Head model. Currently only hallo2 is supported
--result_dirstr/path/to/outputOutput directory (slides, subtitles, videos, etc.)
--paper_latex_rootstr/path/to/latex_projRoot directory of the LaTeX paper project
--ref_imgstr/path/to/ref_img.pngReference image (must be square portrait)
--ref_audiostr/path/to/ref_audio.wavReference audio (recommended: ~10s)
--ref_textstrNoneOptional reference text (for style guidance for subtitles)
--beamer_templete_promptstrNoneOptional reference text (for style guidance for slides)
--gpu_listlist[int]""GPU list for parallel execution (used in cursor generation and Talking Head rendering)
--if_tree_searchboolTrueWhether to enable tree search for slide layout refinement
--stagestr"[0]"Pipeline stages to run (e.g., [0] full pipeline, [1,2,3] partial stages)
--talking_head_envstr/path/to/hallo2_envpython environment path for talking-head generation

📊 Evaluation: Paper2Video

Metrics

Unlike natural video generation, academic presentation videos serve a highly specialized role: they are not merely about visual fidelity but about communicating scholarship. This makes it difficult to directly apply conventional metrics from video synthesis(e.g., FVD, IS, or CLIP-based similarity). Instead, their value lies in how well they disseminate research and amplify scholarly visibility.From this perspective, we argue that a high-quality academic presentation video should be judged along two complementary dimensions:

For the Audience

  • The video is expected to faithfully convey the paper’s core ideas.
  • It should remain accessible to diverse audiences.

For the Author

  • The video should foreground the authors’ intellectual contribution and identity.
  • It should enhance the work’s visibility and impact.

To capture these goals, we introduce evaluation metrics specifically designed for academic presentation videos: Meta Similarity, PresentArena, PresentQuiz, IP Memory.

Run Eval

  • Prepare the environment:
cd src/evaluation
conda create -n p2v_e python=3.10
conda activate p2v_e
pip install -r requirements.txt
  • For MetaSimilarity and PresentArena:
python MetaSim_audio.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python MetaSim_content.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
python PresentArena.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For PresentQuiz, first generate questions from paper and eval using Gemini:
cd PresentQuiz
python create_paper_questions.py ----paper_folder /path/to/data
python PresentQuiz.py --r /path/to/result_dir --g /path/to/gt_dir --s /path/to/save_dir
  • For IP Memory, first generate question pairs from generated videos and eval using Gemini:
cd IPMemory
python construct.py
python ip_qa.py

See the codes for more details!

👉 Paper2Video Benchmark is available at: HuggingFace


😼 Fun: Paper2Video for Paper2Video

Check out How Paper2Video for Paper2Video:

output.mp4

🙏 Acknowledgements

  • The souces of the presentation videos are SlideLive and YouTuBe.
  • We thank all the authors who spend a great effort to create presentation videos!
  • We thank CAMEL for open-source well-organized multi-agent framework codebase.
  • We thank the authors of Hallo2 and Paper2Poster for their open-sourced codes.
  • We thank Wei Jia for his effort in collecting the data and implementing the baselines. We also thank all the participants involved in the human studies.
  • We thank all the Show Lab @ NUS members for support!

📌 Citation

If you find our work useful, please cite:

@misc{paper2video,
title={Paper2Video: Automatic Video Generation from Scientific Papers}, author={Zeyu Zhu and Kevin Qinghong Lin and Mike Zheng Shou},
year={2025},
eprint={2510.05096},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.05096}, }

Star History

About

Automatic Video Generation from Scientific Papers

Resources

Stars

2.4k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages