Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

mlx-quantum-sim

GPU-accelerated quantum circuit simulator for Apple Silicon, with real hardware noise models.

Why I Built This

I had a quantum algorithm idea. Wasn't sure if it worked.

The options were:

  • Rent an A100 on the cloud → $2-5 per experiment, maybe $2,000/year
  • Buy an RTX 4090 → $1,600 upfront
  • Use my MacBook Pro (M1 Max, already on my desk) → $0

I chose the MacBook.

The idea might fail. Six of my first experiments did fail. If I'd rented cloud GPUs for those, I'd have wasted hundreds of dollars discovering dead ends.

Instead, I validated everything locally for $0. When something finally worked, then it made sense to consider scaling up.

This simulator exists because: you shouldn't have to pay to find out if your idea is bad.

What This Is

A quantum circuit simulator that runs on Apple Metal GPU via MLX. Includes calibrated noise profiles from Google Willow, IBM Heron, and QuTech Tuna-9 — extracted from publicly available calibration data.

Honest Comparison: Apple Silicon vs NVIDIA

All qubit counts assume complex64 (float32) precision. Subtract ~1 qubit for complex128.

M1 Max (64GB)M2 Ultra (192GB)RTX 4090 (24GB)A100 cloud (80GB)Colab T4 (free, 16GB)
Max qubits (float32)~32~34~31~33~30
Speed (30q circuit)~5 ms~4 ms~0.3 ms~0.3 ms~1 ms
1000 circuits (30q)~5 sec~4 sec~0.3 sec~0.3 sec~1 sec
Power per experiment0.06 Wh0.07 Wh0.04 WhN/AN/A
Extra hardware cost$0 (if you own a Mac)$5,000$1,600$1.10/hr$0
Setup timepip install mlxSameInstall CUDA + driversCloud setup + SSHOpen browser
Offline / private❌ (cloud)❌ (cloud)
Other usesDaily workDaily workML / gamingML / trainingML (limited)

Where Apple Silicon Wins

  • Zero marginal cost. If you own a Mac, every experiment is free. No cloud bills, no GPU purchase.
  • Large unified memory. M2 Ultra 192GB → 34 qubits. No single consumer GPU matches this.
  • Fast iteration. Edit code → run → results. No uploading, no environment setup, no waiting for cloud instances.
  • Privacy. Circuits never leave your machine.

Where NVIDIA / Cloud Wins

  • Raw speed. 10-15x faster per circuit. For production workloads, this matters.
  • Ecosystem. cuQuantum, qsim, Qiskit Aer CUDA — mature and highly optimized.
  • Scale. Multi-GPU and cluster computing. Apple can't do this.
  • Free tier. Google Colab gives you a T4 GPU for free. Honest alternative for small experiments.

Why Not Just Use Google Colab?

Colab is a strong free option for small experiments. Use mlx-quantum-sim when you need:

  • Offline access (no internet required)
  • More than 16GB GPU memory (Colab T4 limit = ~30 qubits)
  • Long-running experiments (Colab disconnects after ~12 hours)
  • Privacy (proprietary quantum circuits stay local)
  • Reproducibility (no session timeouts or random disconnects)

The Real Decision

"I have a quantum algorithm idea. Should I rent cloud GPUs?"
If you're EXPLORING (not sure it works):
→ Use your Mac. It's free. Most ideas fail. Save your money.
If you've VALIDATED (it works, need to scale):
→ Switch to NVIDIA / cloud. Pay for speed when you know it's worth it.
Apple Silicon = lab notebook (cheap, fast iteration)
NVIDIA = production machine (expensive, maximum throughput)

When to Use This vs NVIDIA

ScenarioUse mlx-quantum-simUse NVIDIA/cloud
Testing a new quantum algorithmOverkill
Homework / courseworkOverkill
Prototyping with noise modelsOverkill
Publishing a paper (small circuits)Optional
Running 100K+ circuits❌ Too slow
Production / deployment
Training quantum ML models⚠️ (small scale OK)✅ (large scale)
33+ qubit simulation✅ (M2 Ultra only)❌ (need multi-GPU)

Features

  • Metal GPU acceleration: All gate operations on Apple Silicon GPU — no CPU round trips
  • Real hardware noise: Google Willow (105 qubits, 364 CZ pairs), IBM Heron, QuTech Tuna-9
  • Stochastic trajectory simulation: Same approach as Google's qsim
  • Up to ~33 qubits on M2 Ultra 192GB (statevector, O(2^n) memory)

Quick Start

pip install mlx
frommlx_quantum_simimportMLXQuantumSimulatorfromnoise_profilesimportWILLOW_NOISE# Ideal simulationsim=MLXQuantumSimulator(3)
sim.h(0)
sim.cx(0, 1)
sim.cx(1, 2)
probs=sim.measure_probs() # GHZ state# With Google Willow noisesim=MLXQuantumSimulator(3, noise_profile=WILLOW_NOISE, n_trajectories=100)
sim.h(0)
sim.cx(0, 1)
probs=sim.measure_probs() # Fidelity < 1 due to realistic noise

Noise Profiles

ProfileSourceCX ErrorT1T2
WILLOW_NOISEcirq-google calibration (willow_pink)0.34%70 μs49 μs
HERON_NOISEIBM published specs0.50%100 μs100 μs
T9_NOISEQuTech published specs1.0%50 μs20 μs

Full per-qubit Willow calibration (T1, T2, readout errors, 364 CZ pairs) in willow_calibration.json.

Noise Channels

  • Depolarizing (single and two-qubit)
  • Amplitude damping (T1 decay)
  • Phase damping (T2 dephasing)
  • Thermal noise (gate-time dependent)
  • Measurement readout error

Gates

H, X, Y, Z, S, T, Rx(θ), Ry(θ), Rz(θ), CNOT/CX, CZ, SWAP

API: sim.ry(theta, qubit) — angle first, qubit second.

Benchmark (M1 Max, 3 qubits, 33 gates)

ModeTime/circuit
Ideal (no noise)2.7 ms
Willow (100 trajectories)495 ms

Mirror Circuit Fidelity

DepthIdealWillowHeronT-9
11.0000.9590.9660.922
51.0000.9420.8650.722
101.0000.8650.8610.490

Ordering matches hardware quality: Willow > Heron > T-9. ✓

Contributing

PRs welcome — especially:

  • CUDA/PyTorch backend (for NVIDIA users)
  • Additional noise profiles (Rigetti, IonQ, Quantinuum)
  • Density matrix simulation
  • Performance optimization

License

Apache License 2.0 (consistent with cirq-google). See LICENSE.

Citation

@software{mlx_quantum_sim,
author = {Huang, Sheng-Kai},
title = {mlx-quantum-sim: GPU-accelerated quantum simulator for Apple Silicon},
year = {2026},
url = {https://github.com/akaiHuang/mlx-quantum-sim}
}

Acknowledgments

Noise calibration data from Google's cirq-google (Apache 2.0). Trajectory-based noise simulation follows the approach of Google's qsim.

About

GPU-accelerated quantum circuit simulator for Apple Silicon (MLX) with Google Willow, IBM Heron, and QuTech Tuna-9 noise models. Up to 34 qubits on M2 Ultra.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages