Skip to content

Repository files navigation

A PyTorch-native and Flexible Inference Engine with
Hybrid Cache Acceleration and Parallelism for 🤗DiTs
Featured|HelloGitHub

BaselineSCM S S*SCM F D*SCM U D*+TS+compile+FP8*
24.85s15.4s11.4s8.2s8.2s🎉7.1s🎉4.5s

Scheme: DBCache + SCM(steps_computation_mask) + TS(TaylorSeer) + FP8*, L20x1, S*: static cache,
D*: dynamic cache, S: Slow, F: Fast, U: Ultra Fast, TS: TaylorSeer, FP8*: FP8 DQ + Sage, FLUX.1-Dev

U*: Ulysses Attention, UAA: Ulysses Anything Attenton, UAA*: UAA + Gloo, Device: NVIDIA L20
FLUX.1-Dev w/o CPU Offload, 28 steps; Qwen-Image w/ CPU Offload, 50 steps; Gloo: Extra All Gather w/ Gloo

CP2 U*CP2 UAA*L20x1CP2 UAA*CP2 U*L20x1CP2 UAA*
FLUX, 13.87s🎉13.88s23.25s🎉13.75sQwen, 132s181s🎉133s
1024x10241024x10241008x10081008x10081312x13121328x13281328x1328
✔️U* ✔️UAA✔️U* ✔️UAANO CP❌U* ✔️UAA✔️U* ✔️UAANO CP❌U* ✔️UAA

SGLang Cache-DiT News

🔥Hightlight

We are excited to announce that the 🎉v1.1.0 version of cache-dit has finally been released! It brings 🔥Context Parallelism and 🔥Tensor Parallelism to cache-dit, thus making it a PyTorch-native and Flexible Inference Engine for 🤗DiTs. Key features: Unified Cache APIs, Forward Pattern Matching, Block Adapter, DBCache, DBPrune, Cache CFG, TaylorSeer, SCM, Context Parallelism (w/ UAA), Tensor Parallelism and 🎉SOTA performance.

pip3 install -U cache-dit # Also, pip3 install git+https://github.com/huggingface/diffusers.git (latest)

You can install the stable release of cache-dit from PyPI, or the latest development version from GitHub. Then try ♥️ Cache Acceleration with just one line of code ~ ♥️

>>>importcache_dit>>>fromdiffusersimportDiffusionPipeline>>>pipe=DiffusionPipeline.from_pretrained("Qwen/Qwen-Image") # Can be any diffusion pipeline>>>cache_dit.enable_cache(pipe) # One-line code with default cache options.>>>output=pipe(...) # Just call the pipe as normal.>>>stats=cache_dit.summary(pipe) # Then, get the summary of cache acceleration stats.>>>cache_dit.disable_cache(pipe) # Disable cache and run original pipe.

📚Core Features

  • 🎉Full 🤗Diffusers Support: Notably, cache-dit now supports nearly all of Diffusers' DiT-based pipelines, include 30+ series, nearly 100+ pipelines, such as FLUX.1, Qwen-Image, Qwen-Image-Lightning, Wan 2.1/2.2, HunyuanImage-2.1, HunyuanVideo, HiDream, AuraFlow, CogView3Plus, CogView4, CogVideoX, LTXVideo, ConsisID, SkyReelsV2, VisualCloze, PixArt, Chroma, Mochi, SD 3.5, DiT-XL, etc.
  • 🎉Extremely Easy to Use: In most cases, you only need one line of code: cache_dit.enable_cache(...). After calling this API, just use the pipeline as normal.
  • 🎉Easy New Model Integration: Features like Unified Cache APIs, Forward Pattern Matching, Automatic Block Adapter, Hybrid Forward Pattern, and Patch Functor make it highly functional and flexible. For example, we achieved 🎉 Day 1 support for HunyuanImage-2.1 with 1.7x speedup w/o precision loss—even before it was available in the Diffusers library.
  • 🎉State-of-the-Art Performance: Compared with algorithms including Δ-DiT, Chipmunk, FORA, DuCa, TaylorSeer and FoCa, cache-dit achieved the SOTA performance w/ 7.4x↑🎉 speedup on ClipScore!
  • 🎉Support for 4/8-Steps Distilled Models: Surprisingly, cache-dit's DBCache works for extremely few-step distilled models—something many other methods fail to do.
  • 🎉Compatibility with Other Optimizations: Designed to work seamlessly with torch.compile, Quantization (torchao, 🔥nunchaku), CPU or Sequential Offloading, 🔥Context Parallelism, 🔥Tensor Parallelism, etc.
  • 🎉Hybrid Cache Acceleration: Now supports hybrid Block-wise Cache + Calibrator schemes (e.g., DBCache or DBPrune + TaylorSeerCalibrator). DBCache or DBPrune acts as the Indicator to decide when to cache, while the Calibrator decides how to cache. More mainstream cache acceleration algorithms (e.g., FoCa) will be supported in the future, along with additional benchmarks—stay tuned for updates!
  • 🤗Diffusers Ecosystem Integration: 🔥cache-dit has joined the Diffusers community ecosystem as the first DiT-specific cache acceleration framework! Check out the documentation here:

The comparison between cache-dit and other algorithms shows that within a speedup ratio (TFLOPs) less than 🎉4x, cache-dit achieved the SOTA performance. Please refer to 📚Benchmarks for more details.

MethodTFLOPs(↓)SpeedUp(↑)ImageReward(↑)Clip Score(↑)
[FLUX.1-dev]: 50 steps3726.871.00×0.989832.404
Chipmunk1505.872.47×0.993632.776
FORA(N=3)1320.072.82×0.977632.266
DBCache(S)1400.082.66×1.006532.838
DuCa(N=5)978.763.80×0.995532.241
TeaCache(l=0.8)892.354.17×0.868331.704
TaylorSeer(N=4,O=2)1042.273.57×0.985732.413
DBCache(S)+TS1153.053.23×1.022132.819
DBCache(M)+TS944.753.94×1.010732.865
FoCa(N=5)893.544.16×1.002932.948
[FLUX.1-dev]: 22% steps818.294.55×0.818331.772
TaylorSeer(N=7,O=2)670.445.54×0.912832.128
FoCa(N=8)596.076.24×0.950232.706
DBCache(F)+TS651.905.72x0.952632.568
DBCache(U)+TS505.477.37x0.864532.719

🎉Surprisingly, cache-dit still works in the extremely few-step distill model, such as Qwen-Image-Lightning, with the F16B16 config, the PSNR is 34.8 and the ImageReward is 1.26. It maintained a relatively high precision.

ConfigPSNR(↑)Clip Score(↑)ImageReward(↑)TFLOPs(↓)SpeedUp(↑)
[Full 4 steps]INF35.57971.2630274.331.00x
F24B2436.324235.62241.2630264.741.04x
F16B1634.816335.61091.2614244.251.12x
F12B1233.895335.65351.2549234.631.17x
F8B833.137435.72841.2517224.291.22x
F1B031.831735.66511.2397206.901.33x

🔥Supported DiTs

Tip

One Model Series may contain many pipelines. cache-dit applies optimizations at the Transformer level; thus, any pipelines that include the supported transformer are already supported by cache-dit. ✔️: known work and official supported now; ✖️: unofficial supported now, but maybe support in the future; Q: 4-bits models w/ nunchaku + SVDQ W4A4; 🔥FLUX.2: 24B + 32B = 56B; 🔥Z-Image: 6B

📚ModelCacheCPTP📚ModelCacheCPTP
🔥Z-Image✔️🔥✔️🔥✔️🔥🔥Ovis-Image✔️🔥✖️✖️
🔥FLUX.2: 56B✔️🔥✔️🔥✔️🔥🔥HuyuanVideo 1.5✔️🔥✖️✖️
🎉FLUX.1✔️✔️✔️🎉FLUX.1 Q✔️✔️✖️
🎉FLUX.1-Fill✔️✔️✔️🎉Qwen-Image Q✔️✔️✖️
🎉Qwen-Image✔️✔️✔️🎉Qwen...Edit Q✔️✔️✖️
🎉Qwen...Edit✔️✔️✔️🎉Qwen...E...Plus Q✔️✔️✖️
🎉Qwen...Lightning✔️✔️✔️🎉Qwen...Light Q✔️✔️✖️
🎉Qwen...Control..✔️✔️✔️🎉Qwen...E...Light Q✔️✔️✖️
🎉Wan 2.1 I2V/T2V✔️✔️✔️🎉Mochi✔️✖️✔️
🎉Wan 2.1 VACE✔️✔️✔️🎉HiDream✔️✖️✖️
🎉Wan 2.2 I2V/T2V✔️✔️✔️🎉HunyunDiT✔️✖️✔️
🎉HunyuanVideo✔️✔️✔️🎉Sana✔️✖️✖️
🎉ChronoEdit✔️✔️✔️🎉Bria✔️✖️✖️
🎉CogVideoX✔️✔️✔️🎉SkyReelsV2✔️✔️✔️
🎉CogVideoX 1.5✔️✔️✔️🎉Lumina 1/2✔️✖️✔️
🎉CogView4✔️✔️✔️🎉DiT-XL✔️✔️✖️
🎉CogView3Plus✔️✔️✔️🎉Allegro✔️✖️✖️
🎉PixArt Sigma✔️✔️✔️🎉Cosmos✔️✖️✖️
🎉PixArt Alpha✔️✔️✔️🎉OmniGen✔️✖️✖️
🎉Chroma-HD✔️✔️️✔️🎉EasyAnimate✔️✖️✖️
🎉VisualCloze✔️✔️✔️🎉StableDiffusion3✔️✖️✖️
🎉HunyuanImage✔️✔️✔️🎉PRX T2I✔️✖️✖️
🎉Kandinsky5✔️✔️️✔️️🎉Amused✔️✖️✖️
🎉LTXVideo✔️✔️✔️🎉AuraFlow✔️✖️✖️
🎉ConsisID✔️✔️✔️🎉LongCatVideo✔️✖️✖️
🔥Click here to show many Image/Video cases🔥

🎉Now, cache-dit covers almost All Diffusers' DiT Pipelines🎉
🔥Qwen-Image | Qwen-Image-Edit | Qwen-Image-Edit-Plus 🔥
🔥FLUX.1 | Qwen-Image-Lightning 4/8 Steps | Wan 2.1 | Wan 2.2 🔥
🔥HunyuanImage-2.1 | HunyuanVideo | HunyuanDiT | HiDream | AuraFlow🔥
🔥CogView3Plus | CogView4 | LTXVideo | CogVideoX | CogVideoX 1.5 | ConsisID🔥
🔥Cosmos | SkyReelsV2 | VisualCloze | OmniGen 1/2 | Lumina 1/2 | PixArt🔥
🔥Chroma | Sana | Allegro | Mochi | SD 3/3.5 | Amused | ... | DiT-XL🔥

🔥Wan2.2 MoE | +cache-dit:2.0x↑🎉 | HunyuanVideo | +cache-dit:2.1x↑🎉

🔥Qwen-Image | +cache-dit:1.8x↑🎉 | FLUX.1-dev | +cache-dit:2.1x↑🎉

🔥Qwen...Lightning | +cache-dit:1.14x↑🎉 | HunyuanImage | +cache-dit:1.7x↑🎉

🔥Qwen-Image-Edit | Input w/o Edit | Baseline | +cache-dit:1.6x↑🎉 | 1.9x↑🎉

🔥FLUX-Kontext-dev | Baseline | +cache-dit:1.3x↑🎉 | 1.7x↑🎉 | 2.0x↑ 🎉

🔥HiDream-I1 | +cache-dit:1.9x↑🎉 | CogView4 | +cache-dit:1.4x↑🎉 | 1.7x↑🎉

🔥CogView3 | +cache-dit:1.5x↑🎉 | 2.0x↑🎉| Chroma1-HD | +cache-dit:1.9x↑🎉

🔥Mochi-1-preview | +cache-dit:1.8x↑🎉 | SkyReelsV2 | +cache-dit:1.6x↑🎉

🔥VisualCloze-512 | Model | Cloth | Baseline | +cache-dit:1.4x↑🎉 | 1.7x↑🎉

🔥LTX-Video-0.9.7 | +cache-dit:1.7x↑🎉 | CogVideoX1.5 | +cache-dit:2.0x↑🎉

🔥OmniGen-v1 | +cache-dit:1.5x↑🎉 | 3.3x↑🎉 | Lumina2 | +cache-dit:1.9x↑🎉

🔥Allegro | +cache-dit:1.36x↑🎉 | AuraFlow-v0.3 | +cache-dit:2.27x↑🎉

🔥Sana | +cache-dit:1.3x↑🎉 | 1.6x↑🎉| PixArt-Sigma | +cache-dit:2.3x↑🎉

🔥PixArt-Alpha | +cache-dit:1.6x↑🎉 | 1.8x↑🎉| SD 3.5 | +cache-dit:2.5x↑🎉

🔥Asumed | +cache-dit:1.1x↑🎉 | 1.2x↑🎉 | DiT-XL-256 | +cache-dit:1.8x↑🎉
♥️ Please consider to leave a ⭐️ Star to support us ~ ♥️

📖Table of Contents

For more advanced features such as Unified Cache APIs, Forward Pattern Matching, Automatic Block Adapter, Hybrid Forward Pattern, Patch Functor, DBCache, DBPrune, TaylorSeer Calibrator, SCM, Hybrid Cache CFG, Context Parallelism (w/ UAA) and Tensor Parallelism, please refer to the 🎉User_Guide.md for details.

👋Contribute

How to contribute? Star ⭐️ this repo to support us or check CONTRIBUTE.md.

Star History Chart

🎉Projects Using CacheDiT

Here is a curated list of open-source projects integrating CacheDiT, including popular repositories like jetson-containers, flux-fast, sdnext and 🔥SGLang Diffusion. 🎉CacheDiT has been recommended by many famous opensource projects: 🔥Z-Image, 🔥Wan 2.2, 🔥Qwen-Image, 🔥LongCat-Video, Qwen-Image-Lightning, Kandinsky-5, LeMiCa, 🤗diffusers, HelloGitHub and GaintPandaCV.

©️Acknowledgements

Special thanks to vipshop's Computer Vision AI Team for supporting document, testing and production-level deployment of this project. We learned the design and reused code from the following projects: 🤗diffusers, ParaAttention, xDiT, TaylorSeer and LeMiCa.

©️Citations

@misc{cache-dit@2025,
title={cache-dit: A PyTorch-native and Flexible Inference Engine with Hybrid Cache Acceleration and Parallelism for DiTs.},
url={https://github.com/vipshop/cache-dit.git},
note={Open-source software available at https://github.com/vipshop/cache-dit.git},
author={DefTruth, vipshop.com},
year={2025}
}

About

A Unified and Flexible Inference Engine with Cache Acceleration, Parallelism and Quantization for 🤗Diffusers.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages