b200
Here are 12 public repositories matching this topic...
Serving and benchmarking 9 frontier MoE checkpoints (490 GiB+) with vLLM on a single 8x B200 node, plus a pre-download HBM fit checker
-
Updated
Aug 18, 2026 - Python
Deploy DeepSeek-V4-Flash-0731 on dual NVIDIA RTX PRO 6000 Blackwell GPUs with vLLM PR #41834 (jasl fork) and DSpark speculative decoding, achieving ~200-227 tok/s in no-overseas-network environments.
-
Updated
Aug 22, 2026
Burst-serving Qwen3.8-27B FP8 on an on-demand RunPod B200, joined to a tailnet as b200 — no public endpoint. Measured: what MTP speculative decoding is worth, what it costs, and what is still unmeasured.
-
Updated
Aug 21, 2026 - HTML
One-B200 DeepSeek V4 Flash deployment, correctness, and reproducible benchmark harness
-
Updated
Aug 5, 2026 - Python
Spheron Network is a decentralized GPU and cloud compute marketplace that aggregates enterprise-grade NVIDIA GPU capacity from certified Tier 3/4 data centers worldwide and exposes it through a single on-demand, per-minute billed interface.
-
Updated
Aug 21, 2026
Production-grade multi-model LLM inference + LoRA fine-tuning on bare-metal Slurm clusters (8× B200). SGLang, vLLM, EAGLE, MTP speculative decoding.
-
Updated
Aug 21, 2026 - Python
Improve this page
Add a description, image, and links to the b200 topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the b200 topic, visit your repo's landing page and select "manage topics."