Skip to content

Repository files navigation

Deepgram Python SDK

Built with FernPyPI versionPython 3.10+MIT License

The official Python SDK for Deepgram's automated speech recognition, text-to-speech, and language understanding APIs. Power your applications with world-class speech and Language AI models.

Documentation

Comprehensive API documentation and guides are available at developers.deepgram.com.

Migrating From Earlier Versions

Installation

Install the Deepgram Python SDK using pip:

pip install deepgram-sdk

Reference

  • API Reference - Complete reference for all SDK methods, parameters, and WebSocket connections

Usage

Quick Start

The Deepgram SDK provides both synchronous and asynchronous clients for all major use cases:

Real-time Speech Recognition (Listen v2)

Our newest and most advanced speech recognition model with contextual turn detection (Reference):

fromdeepgramimportDeepgramClientfromdeepgram.core.eventsimportEventTypeclient=DeepgramClient()
withclient.listen.v2.connect(
model="flux-general-en",
encoding="linear16",
sample_rate=16000
) asconnection:
defon_message(message):
print(f"Received {message.type} event")
connection.on(EventType.OPEN, lambda_: print("Connection opened"))
connection.on(EventType.MESSAGE, on_message)
connection.on(EventType.CLOSE, lambda_: print("Connection closed"))
connection.on(EventType.ERROR, lambdaerror: print(f"Error: {error}"))
# Start listening and send audio dataconnection.start_listening()

File Transcription

Transcribe pre-recorded audio files (API Reference):

fromdeepgramimportDeepgramClientclient=DeepgramClient()
withopen("audio.wav", "rb") asaudio_file:
response=client.listen.v1.media.transcribe_file(
request=audio_file.read(),
model="nova-3"
)
print(response.results.channels[0].alternatives[0].transcript)

Text-to-Speech

Generate natural-sounding speech from text (API Reference):

fromdeepgramimportDeepgramClientclient=DeepgramClient()
response=client.speak.v1.audio.generate(
text="Hello, this is a sample text to speech conversion."
)
# Save the audio filewithopen("output.mp3", "wb") asaudio_file:
audio_file.write(response.stream.getvalue())

Text Analysis

Analyze text for sentiment, topics, and intents (API Reference):

fromdeepgramimportDeepgramClientclient=DeepgramClient()
response=client.read.v1.text.analyze(
request={"text": "Hello, world!"},
language="en",
sentiment=True,
summarize=True,
topics=True,
intents=True
)

Voice Agent (Conversational AI)

Build interactive voice agents (Reference):

fromdeepgramimportDeepgramClientfromdeepgram.agent.v1.typesimport (
AgentV1Settings, AgentV1SettingsAgent,
AgentV1SettingsAgentListen, AgentV1SettingsAgentListenProvider_V1,
AgentV1SettingsAudio, AgentV1SettingsAudioInput,
)
fromdeepgram.types.think_settings_v1importThinkSettingsV1fromdeepgram.types.think_settings_v1providerimportThinkSettingsV1Provider_OpenAifromdeepgram.types.speak_settings_v1importSpeakSettingsV1fromdeepgram.types.speak_settings_v1providerimportSpeakSettingsV1Provider_Deepgramclient=DeepgramClient()
withclient.agent.v1.connect() asagent:
settings=AgentV1Settings(
audio=AgentV1SettingsAudio(
input=AgentV1SettingsAudioInput(encoding="linear16", sample_rate=24000)
),
agent=AgentV1SettingsAgent(
listen=AgentV1SettingsAgentListen(
provider=AgentV1SettingsAgentListenProvider_V1(
type="deepgram", model="nova-3"
)
),
think=ThinkSettingsV1(
provider=ThinkSettingsV1Provider_OpenAi(
type="open_ai", model="gpt-4o-mini"
),
prompt="You are a helpful AI assistant.",
),
speak=SpeakSettingsV1(
provider=SpeakSettingsV1Provider_Deepgram(
type="deepgram", model="aura-2-asteria-en"
)
),
),
)
agent.send_settings(settings)
agent.start_listening()

Complete SDK Reference

For comprehensive documentation of all available methods, parameters, and options:

  • API Reference - Complete reference for all SDK methods including:

    • Listen (Speech-to-Text): File transcription, URL transcription, and media processing
    • Speak (Text-to-Speech): Audio generation and voice synthesis
    • Read (Text Intelligence): Text analysis, sentiment, summarization, and topic detection
    • Manage: Project management, API keys, and usage analytics
    • Auth: Token generation and authentication management
    • WebSocket connections: Listen v1/v2, Speak v1, and Agent v1 real-time streaming

Authentication

The Deepgram SDK supports two authentication methods:

Access Token Authentication

Use access tokens for temporary or scoped access (recommended for client-side applications):

fromdeepgramimportDeepgramClient# Explicit access tokenclient=DeepgramClient(access_token="YOUR_ACCESS_TOKEN")
# Or via environment variable DEEPGRAM_TOKENclient=DeepgramClient()
# Generate access tokens using your API keyauth_client=DeepgramClient(api_key="YOUR_API_KEY")
token_response=auth_client.auth.v1.tokens.grant()
token_client=DeepgramClient(access_token=token_response.access_token)

API Key Authentication

Use your Deepgram API key for server-side applications:

fromdeepgramimportDeepgramClient# Explicit API keyclient=DeepgramClient(api_key="YOUR_API_KEY")
# Or via environment variable DEEPGRAM_API_KEYclient=DeepgramClient()

Environment Variables

The SDK automatically discovers credentials from these environment variables:

  • DEEPGRAM_TOKEN - Your access token (takes precedence)
  • DEEPGRAM_API_KEY - Your Deepgram API key

Precedence: Explicit parameters > Environment variables

Async Client

The SDK provides full async/await support for non-blocking operations:

importasynciofromdeepgramimportAsyncDeepgramClientasyncdefmain():
client=AsyncDeepgramClient()
# Async file transcriptionwithopen("audio.wav", "rb") asaudio_file:
response=awaitclient.listen.v1.media.transcribe_file(
request=audio_file.read(),
model="nova-3"
)
# Async WebSocket connectionasyncwithclient.listen.v2.connect(
model="flux-general-en",
encoding="linear16",
sample_rate=16000
) asconnection:
asyncdefon_message(message):
print(f"Received {message.type} event")
connection.on(EventType.MESSAGE, on_message)
awaitconnection.start_listening()
asyncio.run(main())

Exception Handling

The SDK provides detailed error information for debugging and error handling:

fromdeepgramimportDeepgramClientfromdeepgram.core.api_errorimportApiErrorclient=DeepgramClient()
try:
response=client.listen.v1.media.transcribe_file(
request=audio_data,
model="nova-3"
)
exceptApiErrorase:
print(f"Status Code: {e.status_code}")
print(f"Error Details: {e.body}")
print(f"Request ID: {e.headers.get('x-dg-request-id', 'N/A')}")
exceptExceptionase:
print(f"Unexpected error: {e}")

Advanced Features

Raw Response Access

Access raw HTTP response data including headers:

fromdeepgramimportDeepgramClientclient=DeepgramClient()
response=client.listen.v1.media.with_raw_response.transcribe_file(
request=audio_data,
model="nova-3"
)
print(response.headers) # Access response headersprint(response.data) # Access the response object

Request Configuration

Configure timeouts, retries, and other request options:

fromdeepgramimportDeepgramClient# Global client configurationclient=DeepgramClient(timeout=30.0)
# Per-request configurationresponse=client.listen.v1.media.transcribe_file(
request=audio_data,
model="nova-3",
request_options={
"timeout_in_seconds": 60,
"max_retries": 3
}
)

Custom HTTP Client

Use a custom httpx client for advanced networking features:

importhttpxfromdeepgramimportDeepgramClientclient=DeepgramClient(
httpx_client=httpx.Client(
proxies="http://proxy.example.com",
timeout=httpx.Timeout(30.0)
)
)

Custom Transports

Replace the built-in websockets transport with your own implementation for WebSocket-based APIs (Listen, Speak, Agent). This enables alternative protocols (HTTP/2, SSE), test doubles, or proxied connections.

Any class that implements the right methods can be used as a transport — no inheritance required. Pass your class (or a factory callable) as transport_factory when creating a client.

Sync transports

Implement send(), recv(), __iter__(), and close(), then pass the class to DeepgramClient:

fromdeepgramimportDeepgramClientfromdeepgram.core.eventsimportEventTypeclassMyTransport:
def__init__(self, url: str, headers: dict):
... # establish your connectiondefsend(self, data): ... # send str or bytesdefrecv(self): ... # return next messagedef__iter__(self): ... # yield messages until closeddefclose(self): ... # tear down connectionclient=DeepgramClient(api_key="...", transport_factory=MyTransport)
withclient.listen.v1.connect(model="nova-3") asconnection:
connection.on(EventType.MESSAGE, on_message)
connection.start_listening()

Async transports

Implement async def send(), async def recv(), async def __aiter__(), and async def close(), then use AsyncDeepgramClient:

fromdeepgramimportAsyncDeepgramClientclient=AsyncDeepgramClient(api_key="...", transport_factory=MyAsyncTransport)
asyncwithclient.listen.v1.connect(model="nova-3") asconnection:
connection.on(EventType.MESSAGE, on_message)
awaitconnection.start_listening()

See src/deepgram/transport_interface.py for the full protocol definitions (SyncTransport and AsyncTransport).

SageMaker transport

The deepgram-sagemaker package (source) is a ready-made async transport for running Deepgram models on AWS SageMaker endpoints. It uses HTTP/2 bidirectional streaming under the hood, but exposes the same SDK interface — just install the package and swap in a transport_factory:

pip install deepgram-sagemaker # requires Python 3.12+
fromdeepgramimportAsyncDeepgramClientfromdeepgram_sagemakerimportSageMakerTransportFactoryfactory=SageMakerTransportFactory(
endpoint_name="my-deepgram-endpoint",
region="us-west-2",
)
# SageMaker uses AWS credentials (not Deepgram API keys)client=AsyncDeepgramClient(api_key="unused", transport_factory=factory)
asyncwithclient.listen.v1.connect(model="nova-3") asconnection:
connection.on(EventType.MESSAGE, on_message)
awaitconnection.start_listening()

Note: The SageMaker transport is async-only and requires AsyncDeepgramClient.

See examples/27-transcription-live-sagemaker.py for a complete working example.

Retry Configuration

The SDK automatically retries failed requests with exponential backoff:

# Automatic retries for 408, 429, and 5xx status codesresponse=client.listen.v1.media.transcribe_file(
request=audio_data,
model="nova-3",
request_options={"max_retries": 3}
)

Contributing

We welcome contributions to improve this SDK! However, please note that this library is primarily generated from our API specifications.

Development Setup

  1. Install Poetry (if not already installed):

    curl -sSL https://install.python-poetry.org | python - -y --version 1.5.1
  2. Install dependencies:

    poetry install
  3. Install example dependencies:

    poetry run pip install -r examples/requirements.txt
  4. Run tests:

    poetry run pytest -rP .
  5. Run examples:

    python -u examples/07-transcription-live-websocket.py

Contribution Guidelines

See our CONTRIBUTING guide.

Requirements

  • Python 3.10+
  • See pyproject.toml for full dependency list

Community Code of Conduct

Please see our community code of conduct before contributing to this project.

License

This project is licensed under the MIT License - see the LICENSE file for details.