Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

CommBench
CommBench: can LLMs write correct and efficient GPU communication code?

Blog | Join Slack | Twitter/X | Leaderboard | Quick Start | Contributing

Highlights

  • A benchmark for multi-device GPU communication code. 100+ independently runnable examples spanning P2P, collective, expert-parallel (EP), compute–communication fusion, and utilities, each labeled Easy / Medium / Hard.
  • Curated by GPU communication experts. Every example is hand-written by GPU communication experts or expert-extracted from production codebases including Mscclpp, NCCL, NVSHMEM, DeepEP, ThunderKittens, vLLM, and SGLang.
  • Cheat-resistant evaluation harness. Randomized test inputs, edit-region verification, and hidden build scripts prevent hard-coding and library substitution.
  • Correctness and performance. Generated code is compiled, run, and compared against a hand-written reference on both correctness and speed, with optional multi-round refinement from compile/run feedback.

Today's frontier LLMs write excellent single-device code yet consistently fail on multi-device GPU communication, precisely the code that bottlenecks large-scale LLM training and inference. CommBench measures, and aims to help close, that gap.

How It Works

 empty_*.cu/cpp LLM (GPT, Gemini, Claude, Grok, GLM, Qwen …) generated_*.cu/cpp
┌──────────────┐ ┌──────────┐ ┌──────────────────┐
│ // TODO │ ── prompt ──▶│ Model │── code response──▶│ filled-in code │
│ // TODO │ └──────────┘ └──────────────────┘
└──────────────┘ │
▼
build & run & compare
│
▼
summary.json + plots
  1. Prompt — reads the empty_* template (reference code with core logic stripped to // TODO), builds a completion prompt.
  2. Generate — sends the prompt to an LLM. With --max-rounds > 1, compile/run errors are fed back for self-correction.
  3. Evaluate — compiles and runs both reference and generated code, compares correctness and performance (latency, throughput), and produces CSV, plots, and a summary.json. Supported models: OpenAI GPT, Google Gemini, Anthropic Claude, Grok, DeepSeek, GLM, and Qwen.

Benchmark Categories

CategoryWhat it covers
P2PPoint-to-point transfer between a pair of devices
CollectiveGroup operations across all ranks (AllReduce, AllGather, All-to-All, …)
EPDynamic, non-uniform dispatch/combine traffic for MoE models
FusionKernels that interleave communication with compute (e.g., AllGather+GEMM)
UtilitiesConnection setup, buffer registration, topology queries, GPU–CPU FIFO queues

By source, examples span cuda-runtime, ibverbs, Mscclpp, NCCL, NVSHMEM, DeepEP, the NCCL device API, ThunderKittens, vLLM, and SGLang.

Leaderboard

Sorted by Pass×GM ⭐ — pass rate scaled by geometric-mean code quality on passing examples. See the blog post for full metric definitions and case studies.

RankModelPass×GMPass RatePASS+GoodGM‑SpeedupOpen SourcePrice
🥇gpt-5.50.46757.4%30.7%0.813$1.91
🥈gemini-3.1-pro-preview0.30536.6%25.7%0.832$0.26
🥉claude-opus-4-70.28233.7%20.8%0.836$0.21
4️⃣glm-5.10.28129.7%17.8%0.947$0.63
5️⃣kimi-k2.60.27530.7%18.8%0.895$0.10
6️⃣qwen3.7-max0.26926.7%15.8%1.008$0.03
7️⃣deepseek-v4-pro0.19719.8%12.9%0.995$0.02

Metrics

MetricFormulaWhat it measures
Pass×GMPass Rate × GM‑SpeedupPass rate scaled by geometric-mean code quality on passing examples. Primary ranking metric.
Pass RatePASS / TotalFraction of examples where code compiled, ran, and produced correct results.
PASS+Good(on_compare + better) / TotalFraction of all examples with correct and performant code (within −5% of reference).
GM‑SpeedupGM of per-example speedup scoresGeometric mean of generated-vs-reference performance ratios over passing examples, taken across measured data sizes so each data point contributes equally. Computed only over passing examples, so a model that passes more (often harder) examples with mediocre performance can score lower; we therefore rank by Pass×GM.
PriceAverage cost per example (USD).

Top vs. Bottom Model

A detailed comparison of the highest- and lowest-scoring models, gpt-5.5 (Pass×GM = 0.467) and deepseek-v4-pro (Pass×GM = 0.197).

Difficulty Breakdown

GPT-5.5DeepSeek-V4-Pro

Performance Quality Among PASS Examples

GPT-5.5DeepSeek-V4-Pro

Tag and Library Coverage

GPT-5.5DeepSeek-V4-Pro

Key findings:

  • Even the strongest model passes under 60% of examples and produces performant code on only a third.
  • Every model collapses to near-zero coverage on specialized libraries such as Mscclpp, ThunderKittens, and the NCCL device API. Models hallucinate APIs, misplace synchronization, and ship kernels orders of magnitude slower than reference.
  • Multi-round self-correction helps only on commodity libraries and easier tasks. Giving deepseek-v4-pro 5 rounds raises its pass rate from 15.8% to 41.6%, but unlocks neither Hard examples nor specialized libraries.

Quick Start

# 1. Install dependencies
pip install -r requirements.txt
# 2. Set at least one API keyexport OPENAI_API_KEY="..."# → --model gpt-4oexport GOOGLE_API_KEY="..."# → --model gemini-3-pro-previewexport ANTHROPIC_API_KEY="..."# → --model claude-sonnet-4-5-20250929# 3. Verify GPU compiler
nvcc --version # NVIDIA
hipcc --version # AMD# 4. Run an evaluation
python scripts/generate_eval_one.py example001_gpu_comm_single_process \
--model gpt-4o \
--max-rounds 3

This reads the empty_* template, sends it to the model, compiles and runs the generated code, compares it against the reference, and writes a summary.json with correctness and performance results.

Options

python scripts/generate_eval_one.py <example_name> [options]
Options:
--model MODEL LLM to use (default: gpt-4o)
--max-rounds N Max generation rounds with error feedback (default: 1)
--temperature FLOAT Sampling temperature (default: 0.3)
--datasets-dir PATH Base datasets directory (default: ./datasets)
--no-save Don't save generated code --quiet Suppress detailed output

A single example's build_and_run.py can also be run directly:

python build_and_run.py --source ref_gpu_p2p_comm.cpp # build & run reference
python build_and_run.py --compare ref_gpu_p2p_comm.cpp generated_gpu_p2p_comm.cpp # compare ref vs generated

Contributing

To contribute a new example, please follow the requirements in the Dataset Instructions.

Acknowledgements

We thank Mibura and AMD for sponsoring the testbed for this benchmark.

License

MIT License

About

Can LLMs Write Correct and Efficient GPU Communication Code?

Resources

Stars

63 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages