The Cartesia Python library provides convenient access to the Cartesia API from Python.
Our complete API documentation can be found on docs.cartesia.ai.
pip install cartesiaInstantiate and use the client with the following:
fromcartesiaimportCartesiafromcartesia.ttsimportOutputFormat_Raw, TtsRequestIdSpecifierimportosclient=Cartesia(
api_key=os.getenv("CARTESIA_API_KEY"),
)
client.tts.bytes(
model_id="sonic-2",
transcript="Hello, world!",
voice={
"mode": "id",
"id": "694f9389-aac1-45b6-b726-9d9369183238",
},
language="en",
output_format={
"container": "raw",
"sample_rate": 44100,
"encoding": "pcm_f32le",
},
)The SDK also exports an async client so that you can make non-blocking calls to our API.
importasyncioimportosfromcartesiaimportAsyncCartesiafromcartesia.ttsimportOutputFormat_Raw, TtsRequestIdSpecifierclient=AsyncCartesia(
api_key=os.getenv("CARTESIA_API_KEY"),
)
asyncdefmain() ->None:
asyncforoutputinclient.tts.bytes(
model_id="sonic-2",
transcript="Hello, world!",
voice={"id": "694f9389-aac1-45b6-b726-9d9369183238"},
language="en",
output_format={
"container": "raw",
"sample_rate": 44100,
"encoding": "pcm_f32le",
},
):
print(f"Received chunk of size: {len(output)}")
asyncio.run(main())When the API returns a non-success status code (4xx or 5xx response), a subclass of the following error will be thrown.
fromcartesia.core.api_errorimportApiErrortry:
client.tts.bytes(...)
exceptApiErrorase:
print(e.status_code)
print(e.body)The SDK supports streaming responses as well, returning a generator that you can iterate over with a for ... in ... loop:
fromcartesiaimportCartesiafromcartesia.ttsimportControls, OutputFormat_RawParams, TtsRequestIdSpecifierParamsimportosdefget_tts_chunks():
client=Cartesia(
api_key=os.getenv("CARTESIA_API_KEY"),
)
response=client.tts.sse(
model_id="sonic-2",
transcript="Hello world!",
voice={
"id": "f9836c6e-a0bd-460e-9d3c-f7299fa60f94",
"experimental_controls": {
"speed": "normal",
"emotion": [],
},
},
language="en",
output_format={
"container": "raw",
"encoding": "pcm_f32le",
"sample_rate": 44100,
},
)
audio_chunks= []
forchunkinresponse:
audio_chunks.append(chunk)
returnaudio_chunkschunks=get_tts_chunks()
forchunkinchunks:
print(f"Received chunk of size: {len(chunk.data)}")For the lowest latency in advanced usecases (such as streaming in an LLM-generated transcript and streaming out audio), you should use our websockets client:
fromcartesiaimportCartesiafromcartesia.ttsimportTtsRequestEmbeddingSpecifierParams, OutputFormat_RawParamsimportpyaudioimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
voice_id="a0e99841-438c-4a64-b679-ae501e7d6091"transcript="Hello! Welcome to Cartesia"p=pyaudio.PyAudio()
rate=22050stream=None# Set up the websocket connectionws=client.tts.websocket()
# Generate and stream audio using the websocketforoutputinws.send(
model_id="sonic-2", # see: https://docs.cartesia.ai/getting-started/available-modelstranscript=transcript,
voice={"id": voice_id},
stream=True,
output_format={
"container": "raw",
"encoding": "pcm_f32le",
"sample_rate": rate
},
):
buffer=output.audioifnotstream:
stream=p.open(format=pyaudio.paFloat32, channels=1, rate=rate, output=True)
# Write the audio data to the streamstream.write(buffer)
stream.stop_stream()
stream.close()
p.terminate()
ws.close() # Close the websocket connectionfromcartesiaimportCartesiaimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Load your audio file as byteswithopen("path/to/audio.wav", "rb") asf:
audio_data=f.read()
# Convert to audio chunks (20ms chunks used here for a streaming example)# This chunk size is calculated for 16kHz, 16-bit audio: 16000 * 0.02 * 2 = 640 byteschunk_size=640audio_chunks= [audio_data[i:i+chunk_size] foriinrange(0, len(audio_data), chunk_size)]
# Create websocket connection with endpointing parametersws=client.stt.websocket(
model="ink-whisper", # Model (required)language="en", # Language of your audio (required)encoding="pcm_s16le", # Audio encoding format (required)sample_rate=16000, # Audio sample rate (required)min_volume=0.1, # Volume threshold for voice activity detectionmax_silence_duration_secs=0.4, # Maximum silence duration before endpointing
)
# Send audio chunks (streaming approach)forchunkinaudio_chunks:
ws.send(chunk)
# Finalize and closews.send("finalize")
ws.send("done")
# Receive transcription results with word-level timestampsforresultinws.receive():
ifresult['type'] =='transcript':
print(f"Transcription: {result['text']}")
# Handle word-level timestamps if availableif'words'inresultandresult['words']:
print("Word-level timestamps:")
forword_infoinresult['words']:
word=word_info['word']
start=word_info['start']
end=word_info['end']
print(f" '{word}': {start:.2f}s - {end:.2f}s")
ifresult['is_final']:
print("Final result received")
elifresult['type'] =='done':
breakws.close()For real-time streaming applications, here's a more practical async example that demonstrates concurrent audio processing and result handling:
importasyncioimportosfromcartesiaimportAsyncCartesiaasyncdefstreaming_stt_example():
""" Advanced async STT example for real-time streaming applications. This example simulates streaming audio processing with proper error handling and demonstrates the new endpointing and word timestamp features. """client=AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
try:
# Create websocket connection with voice activity detectionws=awaitclient.stt.websocket(
model="ink-whisper", # Model (required)language="en", # Language of your audio (required)encoding="pcm_s16le", # Audio encoding format (required)sample_rate=16000, # Audio sample rate (required)min_volume=0.15, # Volume threshold for voice activity detectionmax_silence_duration_secs=0.3, # Maximum silence duration before endpointing
)
# Simulate streaming audio data (replace with your audio source)asyncdefaudio_stream():
"""Simulate real-time audio streaming - replace with actual audio capture"""# Load audio file for simulationwithopen("path/to/audio.wav", "rb") asf:
audio_data=f.read()
# Stream in 100ms chunks (realistic for real-time processing)chunk_size=int(16000*0.1*2) # 100ms at 16kHz, 16-bitforiinrange(0, len(audio_data), chunk_size):
chunk=audio_data[i:i+chunk_size]
ifchunk:
yieldchunk# Simulate real-time streaming delayawaitasyncio.sleep(0.1)
# Send audio and receive results concurrentlyasyncdefsend_audio():
"""Send audio chunks to the STT websocket"""try:
asyncforchunkinaudio_stream():
awaitws.send(chunk)
print(f"Sent audio chunk of {len(chunk)} bytes")
# Small delay to simulate realtime applicationsawaitasyncio.sleep(0.02)
# Signal end of audio streamawaitws.send("finalize")
awaitws.send("done")
print("Audio streaming completed")
exceptExceptionase:
print(f"Error sending audio: {e}")
asyncdefreceive_transcripts():
"""Receive and process transcription results with word timestamps"""full_transcript=""all_word_timestamps= []
try:
asyncforresultinws.receive():
ifresult['type'] =='transcript':
text=result['text']
is_final=result['is_final']
# Handle word-level timestampsif'words'inresultandresult['words']:
word_timestamps=result['words']
all_word_timestamps.extend(word_timestamps)
ifis_final:
print("Word-level timestamps:")
forword_infoinword_timestamps:
word=word_info['word']
start=word_info['start']
end=word_info['end']
print(f" '{word}': {start:.2f}s - {end:.2f}s")
ifis_final:
# Final result - this text won't changefull_transcript+=text+" "print(f"FINAL: {text}")
else:
# Partial result - may change as more audio is processedprint(f"PARTIAL: {text}")
elifresult['type'] =='done':
print("Transcription completed")
breakexceptExceptionase:
print(f"Error receiving transcripts: {e}")
returnfull_transcript.strip(), all_word_timestampsprint("Starting streaming STT...")
# Use asyncio.gather to run audio sending and transcript receiving concurrently_, (final_transcript, word_timestamps) =awaitasyncio.gather(
send_audio(),
receive_transcripts()
)
print(f"\nComplete transcript: {final_transcript}")
print(f"Total words with timestamps: {len(word_timestamps)}")
# Clean upawaitws.close()
exceptExceptionase:
print(f"STT streaming error: {e}")
finally:
awaitclient.close()
# Run the exampleif__name__=="__main__":
asyncio.run(streaming_stt_example())For processing pre-recorded audio files, use the batch STT API which supports uploading complete audio files for transcription:
fromcartesiaimportCartesiaimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Transcribe an audio file with word-level timestampswithopen("path/to/audio.wav", "rb") asaudio_file:
response=client.stt.transcribe(
file=audio_file, # Audio file to transcribemodel="ink-whisper", # STT model (required)language="en", # Language of the audio (optional)timestamp_granularities=["word"], # Include word-level timestamps (optional)encoding="pcm_s16le", # Audio encoding (optional)sample_rate=16000, # Audio sample rate (optional)
)
# Access transcription resultsprint(f"Transcribed text: {response.text}")
print(f"Audio duration: {response.duration:.2f} seconds")
# Process word-level timestamps if requestedifresponse.words:
print("\nWord-level timestamps:")
forword_infoinresponse.words:
word=word_info.wordstart=word_info.startend=word_info.endprint(f" '{word}': {start:.2f}s - {end:.2f}s")importasynciofromcartesiaimportAsyncCartesiaimportosasyncdeftranscribe_file():
client=AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
withopen("path/to/audio.wav", "rb") asaudio_file:
response=awaitclient.stt.transcribe(
file=audio_file,
model="ink-whisper",
language="en",
timestamp_granularities=["word"],
)
print(f"Transcribed text: {response.text}")
# Process word timestampsifresponse.words:
forword_infoinresponse.words:
print(f"'{word_info.word}': {word_info.start:.2f}s - {word_info.end:.2f}s")
awaitclient.close()
asyncio.run(transcribe_file())Note: Batch STT also supports OpenAI's audio transcriptions format for easy migration from OpenAI Whisper. See our migration guide for details.
List all available Voices with client.voices.list, which returns an iterable that automatically handles pagination:
fromcartesiaimportCartesiaimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Get all available Voicesvoices=client.voices.list()
forvoiceinvoices:
print(voice)You can also get the complete metadata for a specific Voice, or make a new Voice by cloning from an audio sample:
# Get a specific Voicevoice=client.voices.get(id="a0e99841-438c-4a64-b679-ae501e7d6091")
print("The embedding for", voice.name, "is", voice.embedding)
# Clone a Voice using file datacloned_voice=client.voices.clone(
clip=open("path/to/voice.wav", "rb"),
name="Test cloned voice",
language="en",
mode="similarity", # or "stability"enhance=False, # use enhance=True to clean and denoise the cloning audiodescription="Test voice description"
)importasynciofromcartesiaimportAsyncCartesiaimportosasyncdefmain():
client=AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Connect to the websocketws=awaitclient.tts.websocket()
# Generate speech with timestampsoutput_generate=awaitws.send(
model_id="sonic-2",
transcript="Hello! Welcome to Cartesia's text-to-speech.",
voice={"id": "f9836c6e-a0bd-460e-9d3c-f7299fa60f94"},
output_format={
"container": "raw",
"encoding": "pcm_f32le",
"sample_rate": 44100
},
add_timestamps=True, # Enable word-level timestampsadd_phoneme_timestamps=True, # Enable phonemized timestampsstream=True
)
# Process the streaming response with timestampsall_words= []
all_starts= []
all_ends= []
audio_chunks= []
asyncforoutinoutput_generate:
# Collect audio dataifout.audioisnotNone:
audio_chunks.append(out.audio)
# Process timestamp dataifout.word_timestampsisnotNone:
all_words.extend(out.word_timestamps.words) # List of wordsall_starts.extend(out.word_timestamps.start) # Start time for each word (seconds)all_ends.extend(out.word_timestamps.end) # End time for each word (seconds)awaitws.close()
awaitclient.close()
asyncio.run(main())The SDK is instrumented with automatic retries with exponential backoff. A request will be retried as long as the request is deemed retriable and the number of retry attempts has not grown larger than the configured retry limit (default: 2).
A request is deemed retriable when any of the following HTTP status codes is returned:
Use the max_retries request option to configure this behavior.
client.tts.bytes(..., request_options={
"max_retries": 1
})The SDK defaults to a 60 second timeout. You can configure this with a timeout option at the client or request level.
fromcartesiaimportCartesiaclient=Cartesia(
...,
timeout=20.0,
)
# Override timeout for a specific methodclient.tts.bytes(..., request_options={
"timeout_in_seconds": 1
})# Mix voices togethermixed_voice=client.voices.mix(
voices=[
{"id": "voice_id_1", "weight": 0.25},
{"id": "voice_id_2", "weight": 0.75}
]
)
# Create a new voice from embeddingnew_voice=client.voices.create(
name="Test Voice",
description="Test voice description",
embedding=[...], # List[float] with 192 dimensionslanguage="en"
)You can override the httpx client to customize it for your use-case. Some common use-cases include support for proxies
and transports.
importhttpxfromcartesiaimportCartesiaclient=Cartesia(
...,
httpx_client=httpx.Client(
proxies="http://my.test.proxy.example.com",
transport=httpx.HTTPTransport(local_address="0.0.0.0"),
),
)A full reference for this library is available here.
Note that most of this library is generated programmatically from https://github.com/cartesia-ai/docs — before making edits to a file, verify it's not autogenerated by checking for this comment at the top of the file:
# This file was auto-generated by Fern from our API Definition.
uv pip install -r requirements.txt
uv run pytest -rP -vv tests/custom/test_client.py::test_get_voicesAssuming all your repos are cloned into your home directory:
$ cd~/docs
$ fern generate --group python-sdk --log-level debug --api version-2024-11-13 --preview
$ cd~/cartesia-python
$ git pull ~/docs/fern/apis/version-2024-11-13/.preview/fern-python-sdk
$ git commit --amend -m "manually regenerate from docs"# optionalFrom https://github.com/cartesia-ai/docs click Actions then Release Python SDK. (Requires permissions.)