Python bindings for transcribe.cpp, a C/C++ speech-to-text library built on ggml.
Status: in development. Until wheels are published, use a locally built
libtranscribethrough repo auto-discovery orTRANSCRIBE_LIBRARY.
Upgrading from 0.1? See the
0.2 migration guide,
including the replacement of gpu_device= with exact device objects.
importtranscribe_cppwithtranscribe_cpp.Model("model.gguf") asmodel:
withmodel.session() assession:
result=session.run(pcm_float32_16k_mono)
print(result.text)run() takes mono 16 kHz float32 PCM (buffer-protocol object or sequence). It
does not decode containers or resample; convert audio before calling it.
importnumpyasnppcm=np.asarray(audio, dtype=np.float32) # 1-D, 16 kHz mono# Downmix stereo first; 2-D input is rejected:# pcm = audio.mean(axis=1).astype(np.float32)result=session.run(pcm)Generic run controls use "default" to preserve each model family's shipped
behavior. Models advertising model.supports("pnc") accept pnc="off" or
pnc="on"; models advertising model.supports("itn") accept the equivalent
itn values. The options are available on run(), run_batch(), stream(),
and the one-shot transcribe() helper.
result=session.run(pcm, pnc="off", itn="on")Streaming models expose incremental transcription with committed/tentative
text views — see examples/stream_wav.py:
withmodel.session() assession, session.stream() asstream:
forchunkinpcm_chunks:
stream.feed(chunk)
text=stream.text() # .committed (stable) + .tentativestream.finalize()
result=stream.snapshot() # language, segments, words, tokens, timingsLong transcriptions can be cancelled from another thread with
session.cancel() — the run raises Aborted with the partial transcript on
exc.partial_result (same for OutputTruncated).
Model(backend=...) applies a backend policy ("auto" uses the best
available). transcribe_cpp.backends() returns process-local device objects;
pass one as Model(device=device) for exact selection with no fallback. Persist
a device's device_id, not its runtime handle or index. backend_available(kind)
checks whether a backend policy can currently be satisfied.
device=next(dfordintranscribe_cpp.backends() ifd.device_type=="cpu")
withtranscribe_cpp.Model("model.gguf", device=device) asmodel:
print(model.device)| Variable | Effect |
|---|---|
TRANSCRIBE_BACKEND | overrides the "auto" default; explicit backend= still wins |
TRANSCRIBE_NATIVE_PROVIDER | forces an installed native provider package, for example cu12 |
TRANSCRIBE_LIBRARY | loads exactly this shared library |
Planned wheels will bundle CPU plus platform accelerators;
transcribe-cpp[cu12] will add the CUDA 12 provider.
The binding loads the native library at import and verifies its ABI layout and
version before use. Build a shared library, then run from the repo or point
TRANSCRIBE_LIBRARY at it:
cmake -B build-shared -DTRANSCRIBE_BUILD_SHARED=ON
cmake --build build-shared --target transcribe
cd bindings/python
PYTHONPATH=src uv run --no-project python examples/transcribe_wav.py \
../../models/whisper-tiny.en/whisper-tiny.en-Q5_K_M.gguf ../../samples/jfk.wavNo-model tests always run; model tests skip unless smoke assets are present.
Override paths with TRANSCRIBE_SMOKE_MODEL, TRANSCRIBE_SMOKE_AUDIO, and
TRANSCRIBE_SMOKE_STREAMING_MODEL.
cd bindings/python
TRANSCRIBE_LIBRARY=../../build-shared/src/libtranscribe.dylib \
uv run --extra test pytest- One run/stream at a time per
Modelin 0.x: sessions share the model's compute backend, so serialize runs across sessions (or load one model per worker). See theModeldocstring. - Import package:
transcribe_cpp - Distribution:
transcribe-cpp - License: MIT