Status: pre-alpha / research code. APIs may change without notice; this has not been used in production and hasn't had independent review.
Constrained decoding for LLM inference: at every step, intersect the model's next-token distribution with "which tokens keep generation on a grammar-valid path" -- batched across a serving workload, on the GPU that holds the logits.
Regex, GBNF grammars, and a subset of JSON Schema. The automaton core grew out of a CUDA loopy belief-propagation project: reachability in a constraint automaton is boolean message-passing to a fixpoint, and the (experimental) soft-lookahead pass is its sum-product version.
- Install
- Use it
- How it works
- Why this instead of Outlines / XGrammar / vLLM's built-in guided decoding
- Benchmarks
- Limitations
- Development
- License
Requirements: Python 3.10+, torch >= 2.2, and a C++20 toolchain (the wheel
compiles csrc/ and a small torch op extension at install time). CUDA is
optional -- everything runs CPU-only, including on a laptop; the kernels build
in automatically if nvcc is found.
pip install torch # CPU: --index-url https://download.pytorch.org/whl/cpu
pip install scikit-build-core cmake ninja
pip install --no-build-isolation -e '.[hf]'Regex-constrain a Hugging Face model:
from transformers import AutoModelForCausalLM, AutoTokenizer
from bpdecode.hf import RegexLogitsProcessor
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B")
lp = RegexLogitsProcessor(r"([0-9]{1,3}\.){3}[0-9]{1,3}", tok)
ids = tok("The router IP is ", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, logits_processor=[lp], max_new_tokens=20)[0]))
# ... 192.168.1.100JSON Schema:
from bpdecode.hf import GrammarLogitsProcessor
schema = {
"type": "object",
"properties": {"name": {"type": "string", "maxLength": 24},
"year": {"type": "integer"},
"compiled": {"type": "boolean"}},
"required": ["name", "year", "compiled"],
"additionalProperties": False,
}
gp = GrammarLogitsProcessor.from_json_schema(schema, tok)
# -> {"name": "Rust", "year": 2010, "compiled": true}
# large batch already on the GPU: route through the on-device PDA kernel
# instead of a per-row CPU CFGConstraint (see "How it works" below)
gp_gpu = GrammarLogitsProcessor.from_json_schema(schema, tok, device="cuda")vLLM: from bpdecode.vllm import RegexLogitsProcessorFactory -- pass
SamplingParams(logits_processors=[factory.make(pattern)]).
Batched, on tensors (the serving path -- CPU or CUDA by tensor device):
from bpdecode.batch import GrammarCache, ConstraintBatch
cache = GrammarCache(vocab) # compile-once, shared across requests
batch = ConstraintBatch(cache.get(pattern), capacity=256, device="cuda")
batch.add(request_id) # continuous batching: add / evict / reset
batch.apply_mask(active_ids, logits) # one call masks the whole batch
batch.commit(active_ids, sampled_tokens) # one call advances itSee examples/.
No model, no tensors: RegexConstraint is the plain-Python automaton --
it never touches a tensor at runtime, even though the package install still
needs torch present. Useful for testing or embedding the constraint logic
somewhere else:
from bpdecode import RegexConstraint, Vocabulary
vocab = Vocabulary.from_tokens(["a", "b", "ab", "0", "1", "<eos>"], eos_id=5)
con = RegexConstraint(r"[01]+", vocab)
con.accepts(vocab.token_bytes.index(b"0")) # -> True- Regular grammars (regex, and GBNF/JSON-Schema rules with no recursion)
compile to a byte DFA and a dense
tok_nexttable -- a decode step is a gather. - Context-free grammars run a config-set pushdown automaton
(
grammar.pda); masks are memoised on the compiled grammar (keyed by config-set) and regular sub-loops (json-char*etc.) are spliced out to the dense path. - CUDA:
csrc/has batched kernels (one warp per request) for both paths, validated on an RTX 3090 (scripts/gpu_check.sh).GrammarLogitsProcessorpicks a backend:"cpu"(per-row, memoised -- fastest once warm) or"device"(on-device PDA kernel, one launch for the whole batch, no memo).backend="auto"(default) followsdevice.
Outlines, XGrammar, and llguidance are mature, well-tested libraries doing the same core job (constrained decoding via a Rust/C++ FSM), and for most use cases they're the safer choice today given this project's status above. Two reasons to reach for bpdecode instead:
- it's batched natively --
ConstraintBatchmasks/advances an entire serving batch in one kernel call, with a shared compiled grammar and an adaptive mask cache, rather than one matcher object per request; and once a grammar is warm it's faster per token than any of the three (see below). - the soft-lookahead angle -- steering the model away from valid-but-dead-end tokens instead of only hard-masking, the same idea as "expected future grammaticality" in Grammar-Aligned Decoding -- isn't something Outlines/XGrammar/llguidance ship, though it isn't novel research either (see GAD and Constrained Decoding with Speculative Lookaheads). Details and the (honest, mixed) results are in the Benchmarks section.
bench/RESULTS.md. Against outlines_core, xgrammar, llguidance on a
151k-token vocab:
- regex, per-token mask: ~8 µs on GPU (flat across patterns, beats
outlines_coreon realistic ones); ~90-140 µs on CPU. - JSON Schema, warm: ~1 µs -- an order of magnitude under xgrammar, ~50x under llguidance. The cost is a one-time per-grammar warmup.
- soft lookahead: count-based continuation weighting is a negative
result -- it over-extends and doesn't beat hard masking. A model-weighted
variant fixes it (6/6 complete vs 0/6) at the cost of extra forward passes
per step; shipped for regex and CFG / JSON Schema as
bpdecode.lookahead.generate_model_weighted*(bench/lookahead_model_weighted.py). Full numbers, including GPU, inbench/RESULTS.md.
CUDA kernels validated on an RTX 3090 (scripts/gpu_check.sh): differential
vs the CPU reference + compute-sanitizer, clean.
- Regex: no anchors, backreferences, or lookaround.
- JSON Schema: objects emit properties in schema order -- every string
produced validates, but not every valid property ordering is producible.
minimum/maximumon numbers aren't enforced (not expressible as a grammar). - CFG / pushdown: the config-set is bounded (8 alternative stacks x 32 frames deep) to keep it a fixed size. A pathologically ambiguous or deeply recursive grammar silently saturates rather than raising -- fine for JSON-Schema-scale grammars, not verified beyond that.
- Soft lookahead: the count-based version (
build_lookahead/apply_soft_) is a documented negative result -- see Benchmarks. The model-probability-weighted variant that does work (bpdecode.lookahead.generate_model_weighted*) is its own generation loop, not aLogitsProcessor(model.generate()doesn't expose its KV cache to processors), costs real compute (roughly K times slower decoding), and on GPU the CFG path doesn't scale with batch size the way the regex path does -- seebench/RESULTS.mdfor the numbers and why. - CFG
"device"backend:CFGConstraintBatchhas no state-keyed mask memo (a PDA config-set isn't a cheap hashable key the way a DFA state is), so every step pays a real kernel launch -- the CPU backend's warm memo is still faster for small batches or long-running grammars."device"wins when the batch is large and already living on the GPU.
pip install --no-build-isolation -e '.[dev]'
pytest
ruff check .
cmake -S csrc -B build-cpp && cmake --build build-cpp && ctest --test-dir build-cppMIT