Skip to content

Repository files navigation

python.png

Fish Audio Python SDK

PyPI versionPython VersionPyPI - DownloadscodecovLicense

The official Python library for the Fish Audio API

Documentation:Python SDK Guide | API Reference

Important

Changes to PyPI Versioning

For existing users on Fish Audio Python SDK, please note that the starting version is now 1.0.0. The last version before this was 2025.6.3. You may need to adjust your version constraints accordingly.

The original API in the fish_audio_sdk package has NOT been removed, but you will not receive any updates if you continue using the old versioning scheme.

The simplest fix is to update your dependency to fish-audio-sdk>=1.0.0 to continue receiving updates, or by pinning to a specific version like fish-audio-sdk==1.0.0 when installing via your package manager. There are no changes to the API itself in this transition.

If you're using the legacy fish_audio_sdk and would like to switch to the newer, more robust fishaudio package, see the migration guide to upgrade.

Installation

pip install fish-audio-sdk
# With audio playback utilities
pip install fish-audio-sdk[utils]

Authentication

Get your API key from fish.audio/app/api-keys:

export FISH_API_KEY=your_api_key_here

Or provide directly:

fromfishaudioimportFishAudioclient=FishAudio(api_key="your_api_key")

Quick Start

Synchronous:

fromfishaudioimportFishAudiofromfishaudio.utilsimportplay, saveclient=FishAudio()
# Generate audioaudio=client.tts.convert(text="Hello, world!")
# Play or saveplay(audio)
save(audio, "output.mp3")

Asynchronous:

importasynciofromfishaudioimportAsyncFishAudiofromfishaudio.utilsimportplay, saveasyncdefmain():
client=AsyncFishAudio()
audio=awaitclient.tts.convert(text="Hello, world!")
play(audio)
save(audio, "output.mp3")
asyncio.run(main())

Core Features

Text-to-Speech

Selecting a model:

# Recommended for productionproduction_audio=client.tts.convert(
text="Production speech",
model="s2.1-pro",
)

With custom voice:

# Use a specific voice by IDaudio=client.tts.convert(
text="Custom voice",
reference_id="802e3bc2b27e49c2995d23ef70e6ac89"
)

With speed control:

audio=client.tts.convert(
text="Speaking faster!",
speed=1.5# 1.5x speed
)

Reusable configuration:

fromfishaudio.typesimportTTSConfig, Prosodyconfig=TTSConfig(
prosody=Prosody(speed=1.2, volume=-5),
reference_id="933563129e564b19a115bedd57b7406a",
format="wav",
latency="balanced"
)
# Reuse across generationsaudio1=client.tts.convert(text="First message", config=config)
audio2=client.tts.convert(text="Second message", config=config)

Chunk-by-chunk processing:

# Stream and process chunks as they arriveforchunkinclient.tts.stream(text="Long content..."):
send_to_websocket(chunk)
# Or collect all chunksaudio=client.tts.stream(text="Hello!").collect()

Learn more

Speech-to-Text

# Transcribe audiowithopen("audio.wav", "rb") asf:
result=client.asr.transcribe(audio=f.read(), language="en")
print(result.text)
# Access timestamped segmentsforsegmentinresult.segments:
print(f"[{segment.start:.2f}s - {segment.end:.2f}s] {segment.text}")

Learn more

Real-time Streaming

Stream dynamically generated text for conversational AI and live applications:

Synchronous:

deftext_chunks():
yield"Hello, "yield"this is "yield"streaming!"audio_stream=client.tts.stream_websocket(text_chunks(), latency="balanced")
play(audio_stream)

Asynchronous:

asyncdeftext_chunks():
yield"Hello, "yield"this is "yield"streaming!"audio_stream=awaitclient.tts.stream_websocket(text_chunks(), latency="balanced")
play(audio_stream)

Learn more

Voice Cloning

Instant cloning:

fromfishaudio.typesimportReferenceAudio# Clone voice on-the-flywithopen("reference.wav", "rb") asf:
audio=client.tts.convert(
text="Cloned voice speaking",
references=[ReferenceAudio(
audio=f.read(),
text="Text spoken in reference"
)]
)

Persistent voice models:

# Create voice model for reusewithopen("voice_sample.wav", "rb") asf:
voice=client.voices.create(
title="My Voice",
voices=[f.read()],
description="Custom voice clone"
)
# Use the created modelaudio=client.tts.convert(
text="Using my saved voice",
reference_id=voice.id
)

Learn more

Resource Clients

ResourceDescriptionKey Methods
client.ttsText-to-speechconvert(), stream(), stream_websocket()
client.asrSpeech recognitiontranscribe()
client.voicesVoice managementlist(), get(), create(), update(), delete()
client.accountAccount infoget_credits(), get_package()

Error Handling

fromfishaudio.exceptionsimport (
AuthenticationError,
RateLimitError,
ValidationError,
FishAudioError
)
try:
audio=client.tts.convert(text="Hello!")
exceptAuthenticationError:
print("Invalid API key")
exceptRateLimitError:
print("Rate limit exceeded")
exceptValidationErrorase:
print(f"Invalid request: {e}")
exceptFishAudioErrorase:
print(f"API error: {e}")

Resources

License

This project is licensed under the Apache-2.0 License - see the LICENSE file for details.

About

The official Python library for the Fish Audio API.

Topics

Resources

Contributing

Stars

198 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages