Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
SOTA performance

arXivhuggingface

This is a PyTorch/GPU implementation of the paper "Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation", referred to as VFMTok, which establishes new state-of-the-art performance (gFID: 1.33, gIS: 317.4) for class-to-image generation based on the RAR framework.

VFMTok presents the first experimental evidence that features from existing pre-trained vision foundation models (including DINOv2, SigLIP, SigLIP2, etc.) can be directly utilized to reconstruct original images. To accomplish this, VFMTok introduces two innovative components: (1) a region-adaptive quantization framework that minimizes redundancy in pre-trained features on standard 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer's outputs with the foundation model's representations to maintain semantic fidelity.

When integrated into the AR generative models, the trained VFMTok achieves remarkable performance in class-to-image generation while tripling the convergence speed. Additionally, it enables high-fidelity class-conditional synthesis without requiring classifier-free guidance (CFG).

This repo contains:

  • 🪐 A simple PyTorch implementation of VFMTok and various new *state-of-the-art generative models.
  • ⚡️ Pre-trained tokenizer: VFMTok and AR generative models trained on ImageNet.
  • 🛸 Training and evaluation scripts for tokenizer and generative models, which were also provided in here.
  • 🎉 Hugging Face for easy access to pre-trained models.

Release

  • [2025/07/11] 🔥 VFMTok has been released. Checkout the paper for details.🔥
  • [2025/09/18] 🔥 VFMTok has been accepted by NeurIPS 2025! 🔥
  • [2025/10/11] 🔥 Image tokenizers and AR models for class-conditional image generation are released. 🔥
  • [2025/10/11] 🔥 All codes of VFMTok have been released. 🔥

Contents

Install

If you are not using Linux, do NOT proceed.

  1. Clone this repository and navigate to VFMTok-RAR folder
git clone https://github.com/CVMI-Lab/VFMTok-RAR.git
cd VFMTok-RAR
  1. Create the vfmtok environment
conda create -n vfmtok python=3.10 -y
conda activate vfmtok
  1. Install deformable attention module
cd vfmtok/modules/ops
bash make.sh

Model Zoom

In this repo, we release:

  • One image tokenizers: VFMTok(DINOv2).
  • State-of-the-art class-conditional autoregressive generative models ranging from 461M to 1.5B parameters.

1. VQ-VAE models

In this repo, we release one image tokenizer: VFMTok(DINOv2). It directly utilizes the features from the frozen pre-trained VFM -- DINOv2, to reconstruct the image. Besides, VFMToks also designs 2 key components: region-adaptive quantization and semantic reconstruction to reduce the redundancy in the pretrained features and maintain the semantic fidelity, respectively.

MethodtokensrFID (256x256)rIS (256x256)weight
VFMTok2560.89216.2vfmtok-tokenizer.pt

2. AR generation models with classifier-free guidance (CFG).

Once the trained VFMTok(DINOv2) is integrated into autoregressive (AR) generative model -- RAR, it ahieves new state-of-the-art image generation performance. Here we provide 2 types of AR generative models: ultra and vanilla. The ultra AR generative model achieves a new state-of-the-art image synthesis performance, while the vanilla ones also produce significant generation performance.

MethodparamsepochsFIDsFIDISPre.Rec.
RAR-L-ultra461M4001.335.72317.40.780.65
RAR-L-vanilla461M4001.446.03312.80.780.66
RAR-XL-vanilla955M4001.385.86310.20.780.65
RAR-XXL-vanilla1.5B4001.365.86301.30.780.66

3. AR generation without CFG (CFG-free image generation).

The trained VFMTok(DINOv2), when integrated into the AR generation models, can also achieve impressive image generation quality without CFG-guidance (CFG-free guidance).

MethodparamsepochsFIDsFIDISPre.Rec.
VFMTok-L-ultra461M4002.015.34211.10.780.63
RAR-L-vanilla461M4002.025.51210.40.790.63
RAR-XL-vanilla955M4001.745.33233.00.800.63
RAR-XXL-vanilla1.5B4001.655.55253.70.800.63

Training

1. Preparation

  1. Download the DINOv2-L pre-trained foundation model from the official model zoo.
  2. Create symbolic links that point from the locations of the pretrained DINOv2-L model and the ImageNet training dataset folders to this directory.
  3. Create dataset script for your own dataset. Here, we provide a template for training tokenizers and AR generative models using the ImageNet dataset in LMDB format.
ln -s DINOv2-L_folder init_models
ln -s ImageNetFolder imagenet

2.VFMTok Training

  1. Training VFMTok(DINOv2) tokenizer (see scripts/tokenizer/train_tok.sh):
export NODE_COUNT=1
export NODE_RANK=0
export PROC_PER_NODE=8
scripts/autoregressive/torchrun.sh vq_train.py --image-size 336 --results-dir output --mixed-precision none --codebook-slots-embed-dim 12 \
--data-path imagenet/lmdb/train_lmdb --global-batch-size 8 --num-workers 4 --ckpt-every 5000 --epochs 50 \
--transformer-config configs/vit_transformer.yaml --log-every 1 --lr 1e-4 --ema --z-channels 512 \

3. AR generative model training

  1. Training AR generative models (see scripts/autoregressive/run_train.sh)
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4
  1. Resume from an AR generative checkpoint
config_file='configs/training/generator/rar.yaml'
accelerate launch --config_file $1 train_rar.py --config-file ${config_file} --image-size 336 --anno-file imagenet/lmdb/train_lmdb --num-workers 4

4. Evaluation (ImageNet 256x256)

  1. Evaluated a pretrained tokenizer (see scripts/tokenizer/run_tok.sh):
scripts/autoregressive/torchrun.sh vqgan_test.py --vq-model VQ-16 --image-size 336 --output_dir recons --batch-size $1 \
--z-channels 512 --vq-ckpt tokenizer/vfmtok-tokenizer.pt --codebook-slots-embed-dim 12
  1. Evaluate a pretrained AR generative model (see scripts/autoregressive/run_test.sh)
config_file='configs/training/generator/rar.yaml'
iters="checkpoint-$(printf "%06d""$1")"
scripts/autoregressive/torchrun.sh test_net.py --config-file ${config_file} --compile \
--gpt-ckpt snapshot/RAR-L/${iters}/model.safetensors --image-size 256 --image-size-eval 256 --per-proc-batch-size $2 \
--guidance-scale $3 --sample-dir samples --guidance-scale-pow 1

Citation

If you find VFMTok useful for your research and applications, please kindly cite using this BibTeX:

@article{zheng2025vision,
title={Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation},
author={Zheng, Anlin and Wen, Xin and Zhang, Xuanyang and Ma, Chuofan and Wang, Tiancai and Yu, Gang and Zhang, Xiangyu and Qi, Xiaojuan},
journal={arXiv preprint arXiv:2507.08441},
year={2025}
}

License

The majority of this project is licensed under Apacha 2.0 License. Portions of the project are available under separate license of referred projects, detailed in corresponding files.

Acknowledgement

Our codebase builds upon several excellent open-source projects, including LlamaGen, Deformable DETR, VFMTok, RAR and AliTok. We are grateful to the communities behind them.

Contact

This codebase has been cleaned up but has not undergone extensive testing. If you encounter any issues or have questions, please open a GitHub issue. We appreciate your feedback!

About

(NeurIPS 2025, SOTA) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages