Skip to content

Repository files navigation

Cartesia Python Library

fern shieldpypi

The Cartesia Python library provides convenient access to the Cartesia API from Python.

Documentation

Our complete API documentation can be found on docs.cartesia.ai.

Installation

pip install cartesia

Usage

Instantiate and use the client with the following:

fromcartesiaimportCartesiafromcartesia.ttsimportOutputFormat_Raw, TtsRequestIdSpecifierimportosclient=Cartesia(
api_key=os.getenv("CARTESIA_API_KEY"),
)
client.tts.bytes(
model_id="sonic-2",
transcript="Hello, world!",
voice={
"mode": "id",
"id": "694f9389-aac1-45b6-b726-9d9369183238",
},
language="en",
output_format={
"container": "raw",
"sample_rate": 44100,
"encoding": "pcm_f32le",
},
)

Async Client

The SDK also exports an async client so that you can make non-blocking calls to our API.

importasyncioimportosfromcartesiaimportAsyncCartesiafromcartesia.ttsimportOutputFormat_Raw, TtsRequestIdSpecifierclient=AsyncCartesia(
api_key=os.getenv("CARTESIA_API_KEY"),
)
asyncdefmain() ->None:
asyncforoutputinclient.tts.bytes(
model_id="sonic-2",
transcript="Hello, world!",
voice={"id": "694f9389-aac1-45b6-b726-9d9369183238"},
language="en",
output_format={
"container": "raw",
"sample_rate": 44100,
"encoding": "pcm_f32le",
},
):
print(f"Received chunk of size: {len(output)}")
asyncio.run(main())

Exception Handling

When the API returns a non-success status code (4xx or 5xx response), a subclass of the following error will be thrown.

fromcartesia.core.api_errorimportApiErrortry:
client.tts.bytes(...)
exceptApiErrorase:
print(e.status_code)
print(e.body)

Streaming

The SDK supports streaming responses as well, returning a generator that you can iterate over with a for ... in ... loop:

fromcartesiaimportCartesiafromcartesia.ttsimportControls, OutputFormat_RawParams, TtsRequestIdSpecifierParamsimportosdefget_tts_chunks():
client=Cartesia(
api_key=os.getenv("CARTESIA_API_KEY"),
)
response=client.tts.sse(
model_id="sonic-2",
transcript="Hello world!",
voice={
"id": "f9836c6e-a0bd-460e-9d3c-f7299fa60f94",
"experimental_controls": {
"speed": "normal",
"emotion": [],
},
},
language="en",
output_format={
"container": "raw",
"encoding": "pcm_f32le",
"sample_rate": 44100,
},
)
audio_chunks= []
forchunkinresponse:
audio_chunks.append(chunk)
returnaudio_chunkschunks=get_tts_chunks()
forchunkinchunks:
print(f"Received chunk of size: {len(chunk.data)}")

WebSockets

For the lowest latency in advanced usecases (such as streaming in an LLM-generated transcript and streaming out audio), you should use our websockets client:

fromcartesiaimportCartesiafromcartesia.ttsimportTtsRequestEmbeddingSpecifierParams, OutputFormat_RawParamsimportpyaudioimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
voice_id="a0e99841-438c-4a64-b679-ae501e7d6091"transcript="Hello! Welcome to Cartesia"p=pyaudio.PyAudio()
rate=22050stream=None# Set up the websocket connectionws=client.tts.websocket()
# Generate and stream audio using the websocketforoutputinws.send(
model_id="sonic-2", # see: https://docs.cartesia.ai/getting-started/available-modelstranscript=transcript,
voice={"id": voice_id},
stream=True,
output_format={
"container": "raw",
"encoding": "pcm_f32le",
"sample_rate": rate
},
):
buffer=output.audioifnotstream:
stream=p.open(format=pyaudio.paFloat32, channels=1, rate=rate, output=True)
# Write the audio data to the streamstream.write(buffer)
stream.stop_stream()
stream.close()
p.terminate()
ws.close() # Close the websocket connection

Speech-to-Text (STT) with Websockets

fromcartesiaimportCartesiaimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Load your audio file as byteswithopen("path/to/audio.wav", "rb") asf:
audio_data=f.read()
# Convert to audio chunks (20ms chunks used here for a streaming example)# This chunk size is calculated for 16kHz, 16-bit audio: 16000 * 0.02 * 2 = 640 byteschunk_size=640audio_chunks= [audio_data[i:i+chunk_size] foriinrange(0, len(audio_data), chunk_size)]
# Create websocket connection with endpointing parametersws=client.stt.websocket(
model="ink-whisper", # Model (required)language="en", # Language of your audio (required)encoding="pcm_s16le", # Audio encoding format (required)sample_rate=16000, # Audio sample rate (required)min_volume=0.1, # Volume threshold for voice activity detectionmax_silence_duration_secs=0.4, # Maximum silence duration before endpointing
)
# Send audio chunks (streaming approach)forchunkinaudio_chunks:
ws.send(chunk)
# Finalize and closews.send("finalize")
ws.send("done")
# Receive transcription results with word-level timestampsforresultinws.receive():
ifresult['type'] =='transcript':
print(f"Transcription: {result['text']}")
# Handle word-level timestamps if availableif'words'inresultandresult['words']:
print("Word-level timestamps:")
forword_infoinresult['words']:
word=word_info['word']
start=word_info['start']
end=word_info['end']
print(f" '{word}': {start:.2f}s - {end:.2f}s")
ifresult['is_final']:
print("Final result received")
elifresult['type'] =='done':
breakws.close()

Async Streaming Speech-to-Text (STT) with Websockets

For real-time streaming applications, here's a more practical async example that demonstrates concurrent audio processing and result handling:

importasyncioimportosfromcartesiaimportAsyncCartesiaasyncdefstreaming_stt_example():
""" Advanced async STT example for real-time streaming applications. This example simulates streaming audio processing with proper error handling and demonstrates the new endpointing and word timestamp features. """client=AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
try:
# Create websocket connection with voice activity detectionws=awaitclient.stt.websocket(
model="ink-whisper", # Model (required)language="en", # Language of your audio (required)encoding="pcm_s16le", # Audio encoding format (required)sample_rate=16000, # Audio sample rate (required)min_volume=0.15, # Volume threshold for voice activity detectionmax_silence_duration_secs=0.3, # Maximum silence duration before endpointing
)
# Simulate streaming audio data (replace with your audio source)asyncdefaudio_stream():
"""Simulate real-time audio streaming - replace with actual audio capture"""# Load audio file for simulationwithopen("path/to/audio.wav", "rb") asf:
audio_data=f.read()
# Stream in 100ms chunks (realistic for real-time processing)chunk_size=int(16000*0.1*2) # 100ms at 16kHz, 16-bitforiinrange(0, len(audio_data), chunk_size):
chunk=audio_data[i:i+chunk_size]
ifchunk:
yieldchunk# Simulate real-time streaming delayawaitasyncio.sleep(0.1)
# Send audio and receive results concurrentlyasyncdefsend_audio():
"""Send audio chunks to the STT websocket"""try:
asyncforchunkinaudio_stream():
awaitws.send(chunk)
print(f"Sent audio chunk of {len(chunk)} bytes")
# Small delay to simulate realtime applicationsawaitasyncio.sleep(0.02)
# Signal end of audio streamawaitws.send("finalize")
awaitws.send("done")
print("Audio streaming completed")
exceptExceptionase:
print(f"Error sending audio: {e}")
asyncdefreceive_transcripts():
"""Receive and process transcription results with word timestamps"""full_transcript=""all_word_timestamps= []
try:
asyncforresultinws.receive():
ifresult['type'] =='transcript':
text=result['text']
is_final=result['is_final']
# Handle word-level timestampsif'words'inresultandresult['words']:
word_timestamps=result['words']
all_word_timestamps.extend(word_timestamps)
ifis_final:
print("Word-level timestamps:")
forword_infoinword_timestamps:
word=word_info['word']
start=word_info['start']
end=word_info['end']
print(f" '{word}': {start:.2f}s - {end:.2f}s")
ifis_final:
# Final result - this text won't changefull_transcript+=text+" "print(f"FINAL: {text}")
else:
# Partial result - may change as more audio is processedprint(f"PARTIAL: {text}")
elifresult['type'] =='done':
print("Transcription completed")
breakexceptExceptionase:
print(f"Error receiving transcripts: {e}")
returnfull_transcript.strip(), all_word_timestampsprint("Starting streaming STT...")
# Use asyncio.gather to run audio sending and transcript receiving concurrently_, (final_transcript, word_timestamps) =awaitasyncio.gather(
send_audio(),
receive_transcripts()
)
print(f"\nComplete transcript: {final_transcript}")
print(f"Total words with timestamps: {len(word_timestamps)}")
# Clean upawaitws.close()
exceptExceptionase:
print(f"STT streaming error: {e}")
finally:
awaitclient.close()
# Run the exampleif__name__=="__main__":
asyncio.run(streaming_stt_example())

Batch Speech-to-Text (STT)

For processing pre-recorded audio files, use the batch STT API which supports uploading complete audio files for transcription:

fromcartesiaimportCartesiaimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Transcribe an audio file with word-level timestampswithopen("path/to/audio.wav", "rb") asaudio_file:
response=client.stt.transcribe(
file=audio_file, # Audio file to transcribemodel="ink-whisper", # STT model (required)language="en", # Language of the audio (optional)timestamp_granularities=["word"], # Include word-level timestamps (optional)encoding="pcm_s16le", # Audio encoding (optional)sample_rate=16000, # Audio sample rate (optional)
)
# Access transcription resultsprint(f"Transcribed text: {response.text}")
print(f"Audio duration: {response.duration:.2f} seconds")
# Process word-level timestamps if requestedifresponse.words:
print("\nWord-level timestamps:")
forword_infoinresponse.words:
word=word_info.wordstart=word_info.startend=word_info.endprint(f" '{word}': {start:.2f}s - {end:.2f}s")

Async Batch STT

importasynciofromcartesiaimportAsyncCartesiaimportosasyncdeftranscribe_file():
client=AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
withopen("path/to/audio.wav", "rb") asaudio_file:
response=awaitclient.stt.transcribe(
file=audio_file,
model="ink-whisper",
language="en",
timestamp_granularities=["word"],
)
print(f"Transcribed text: {response.text}")
# Process word timestampsifresponse.words:
forword_infoinresponse.words:
print(f"'{word_info.word}': {word_info.start:.2f}s - {word_info.end:.2f}s")
awaitclient.close()
asyncio.run(transcribe_file())

Note: Batch STT also supports OpenAI's audio transcriptions format for easy migration from OpenAI Whisper. See our migration guide for details.

Voices

List all available Voices with client.voices.list, which returns an iterable that automatically handles pagination:

fromcartesiaimportCartesiaimportosclient=Cartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Get all available Voicesvoices=client.voices.list()
forvoiceinvoices:
print(voice)

You can also get the complete metadata for a specific Voice, or make a new Voice by cloning from an audio sample:

# Get a specific Voicevoice=client.voices.get(id="a0e99841-438c-4a64-b679-ae501e7d6091")
print("The embedding for", voice.name, "is", voice.embedding)
# Clone a Voice using file datacloned_voice=client.voices.clone(
clip=open("path/to/voice.wav", "rb"),
name="Test cloned voice",
language="en",
mode="similarity", # or "stability"enhance=False, # use enhance=True to clean and denoise the cloning audiodescription="Test voice description"
)

Requesting Timestamps

importasynciofromcartesiaimportAsyncCartesiaimportosasyncdefmain():
client=AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Connect to the websocketws=awaitclient.tts.websocket()
# Generate speech with timestampsoutput_generate=awaitws.send(
model_id="sonic-2",
transcript="Hello! Welcome to Cartesia's text-to-speech.",
voice={"id": "f9836c6e-a0bd-460e-9d3c-f7299fa60f94"},
output_format={
"container": "raw",
"encoding": "pcm_f32le",
"sample_rate": 44100
},
add_timestamps=True, # Enable word-level timestampsadd_phoneme_timestamps=True, # Enable phonemized timestampsstream=True
)
# Process the streaming response with timestampsall_words= []
all_starts= []
all_ends= []
audio_chunks= []
asyncforoutinoutput_generate:
# Collect audio dataifout.audioisnotNone:
audio_chunks.append(out.audio)
# Process timestamp dataifout.word_timestampsisnotNone:
all_words.extend(out.word_timestamps.words) # List of wordsall_starts.extend(out.word_timestamps.start) # Start time for each word (seconds)all_ends.extend(out.word_timestamps.end) # End time for each word (seconds)awaitws.close()
awaitclient.close()
asyncio.run(main())

Advanced

Retries

The SDK is instrumented with automatic retries with exponential backoff. A request will be retried as long as the request is deemed retriable and the number of retry attempts has not grown larger than the configured retry limit (default: 2).

A request is deemed retriable when any of the following HTTP status codes is returned:

  • 408 (Timeout)
  • 429 (Too Many Requests)
  • 5XX (Internal Server Errors)

Use the max_retries request option to configure this behavior.

client.tts.bytes(..., request_options={
"max_retries": 1
})

Timeouts

The SDK defaults to a 60 second timeout. You can configure this with a timeout option at the client or request level.

fromcartesiaimportCartesiaclient=Cartesia(
...,
timeout=20.0,
)
# Override timeout for a specific methodclient.tts.bytes(..., request_options={
"timeout_in_seconds": 1
})

Mixing voices and creating from embeddings

# Mix voices togethermixed_voice=client.voices.mix(
voices=[
{"id": "voice_id_1", "weight": 0.25},
{"id": "voice_id_2", "weight": 0.75}
]
)
# Create a new voice from embeddingnew_voice=client.voices.create(
name="Test Voice",
description="Test voice description",
embedding=[...], # List[float] with 192 dimensionslanguage="en"
)

Custom Client

You can override the httpx client to customize it for your use-case. Some common use-cases include support for proxies and transports.

importhttpxfromcartesiaimportCartesiaclient=Cartesia(
...,
httpx_client=httpx.Client(
proxies="http://my.test.proxy.example.com",
transport=httpx.HTTPTransport(local_address="0.0.0.0"),
),
)

Reference

A full reference for this library is available here.

Contributing

Note that most of this library is generated programmatically from https://github.com/cartesia-ai/docs — before making edits to a file, verify it's not autogenerated by checking for this comment at the top of the file:

# This file was auto-generated by Fern from our API Definition.

Running tests

uv pip install -r requirements.txt
uv run pytest -rP -vv tests/custom/test_client.py::test_get_voices

Manually generating SDK code from docs

Assuming all your repos are cloned into your home directory:

$ cd~/docs
$ fern generate --group python-sdk --log-level debug --api version-2024-11-13 --preview
$ cd~/cartesia-python
$ git pull ~/docs/fern/apis/version-2024-11-13/.preview/fern-python-sdk
$ git commit --amend -m "manually regenerate from docs"# optional

Automatically generating new SDK releases

From https://github.com/cartesia-ai/docs click Actions then Release Python SDK. (Requires permissions.)

About

The official Cartesia client for Python.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages