Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

The LiveKit icon, the name of the repository and some sample code in the background.

PyPI - VersionPyPI DownloadsSlack communityTwitter FollowAsk DeepWiki for understanding the codebaseLicense


Looking for the JS/TS library? Check out AgentsJS

What is Agents?

The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand.

Features

  • Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case.
  • Integrated job scheduling: Built-in task scheduling and distribution with dispatch APIs to connect end users to agents.
  • Extensive WebRTC clients: Build client applications using LiveKit's open-source SDK ecosystem, supporting all major platforms.
  • Telephony integration: Works seamlessly with LiveKit's telephony stack, allowing your agent to make calls to or receive calls from phones.
  • Exchange data with clients: Use RPCs and other Data APIs to seamlessly exchange data with clients.
  • Semantic turn detection: Uses a transformer model to detect when a user is done with their turn, helps to reduce interruptions.
  • MCP support: Native support for MCP. Integrate tools provided by MCP servers with one loc.
  • Builtin test framework: Write tests and use judges to ensure your agent is performing as expected.
  • Open-source: Fully open-source, allowing you to run the entire stack on your own servers, including LiveKit server, one of the most widely used WebRTC media servers.

Installation

To install the core Agents library, along with plugins for popular model providers:

pip install "livekit-agents[openai,silero,deepgram,cartesia,turn-detector]~=1.0"

Docs and guides

Documentation on the framework and how to use it can be found here

Core concepts

  • Agent: An LLM-based application with defined instructions.
  • AgentSession: A container for agents that manages interactions with end users.
  • entrypoint: The starting point for an interactive session, similar to a request handler in a web server.
  • AgentServer: The main process that coordinates job scheduling and launches agents for user sessions.

Usage

Simple voice agent


fromlivekit.agentsimport (
Agent,
AgentServer,
AgentSession,
JobContext,
RunContext,
cli,
function_tool,
inference,
)
fromlivekit.pluginsimportsilero@function_toolasyncdeflookup_weather(
context: RunContext,
location: str,
):
"""Used to look up weather information."""return {"weather": "sunny", "temperature": 70}
server=AgentServer()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
session=AgentSession(
vad=silero.VAD.load(),
# any combination of STT, LLM, TTS, or realtime API can be used# this example shows LiveKit Inference, a unified API to access different models via LiveKit Cloud# to use model provider keys directly, replace with the following:# from livekit.plugins import deepgram, openai, cartesia# stt=deepgram.STT(model="nova-3"),# llm=openai.LLM(model="gpt-4.1-mini"),# tts=cartesia.TTS(model="sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),stt=inference.STT("deepgram/nova-3", language="multi"),
llm=inference.LLM("openai/gpt-4.1-mini"),
tts=inference.TTS("cartesia/sonic-3", voice="9626c31c-bec5-4cca-baa8-f8ba9e84c8bc"),
)
agent=Agent(
instructions="You are a friendly voice assistant built by LiveKit.",
tools=[lookup_weather],
)
awaitsession.start(agent=agent, room=ctx.room)
awaitsession.generate_reply(instructions="greet the user and ask about their day")
if__name__=="__main__":
cli.run_app(server)

You'll need the following environment variables for this example:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

Multi-agent handoff


This code snippet is abbreviated. For the full example, see multi_agent.py

...
classIntroAgent(Agent):
def__init__(self) ->None:
super().__init__(
instructions=f"You are a story teller. Your goal is to gather a few pieces of information from the user to make the story personalized and engaging.""Ask the user for their name and where they are from"
)
asyncdefon_enter(self):
self.session.generate_reply(instructions="greet the user and gather information")
@function_toolasyncdefinformation_gathered(
self,
context: RunContext,
name: str,
location: str,
):
"""Called when the user has provided the information needed to make the story personalized and engaging. Args: name: The name of the user location: The location of the user """context.userdata.name=namecontext.userdata.location=locationstory_agent=StoryAgent(name, location)
returnstory_agent, "Let's start the story!"classStoryAgent(Agent):
def__init__(self, name: str, location: str) ->None:
super().__init__(
instructions=f"You are a storyteller. Use the user's information in order to make the story personalized."f"The user's name is {name}, from {location}"# override the default model, switching to Realtime API from standard LLMsllm=openai.realtime.RealtimeModel(voice="echo"),
chat_ctx=chat_ctx,
)
asyncdefon_enter(self):
self.session.generate_reply()
@server.rtc_session()asyncdefentrypoint(ctx: JobContext):
userdata=StoryData()
session=AgentSession[StoryData](
vad=silero.VAD.load(),
stt="deepgram/nova-3",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
userdata=userdata,
)
awaitsession.start(
agent=IntroAgent(),
room=ctx.room,
)
...

Testing

Automated tests are essential for building reliable agents, especially with the non-deterministic behavior of LLMs. LiveKit Agents include native test integration to help you create dependable agents.

@pytest.mark.asyncioasyncdeftest_no_availability() ->None:
llm=google.LLM()
asyncAgentSession(llm=llm) assess:
awaitsess.start(MyAgent())
result=awaitsess.run(
user_input="Hello, I need to place an order."
)
result.expect.skip_next_event_if(type="message", role="assistant")
result.expect.next_event().is_function_call(name="start_order")
result.expect.next_event().is_function_call_output()
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="assistant should be asking the user what they would like")
)

Examples

πŸŽ™οΈ Starter Agent

A starter agent optimized for voice conversations.

Code

πŸ”„ Multi-user push to talk

Responds to multiple users in the room via push-to-talk.

Code

🎡 Background audio

Background ambient and thinking audio to improve realism.

Code

πŸ› οΈ Dynamic tool creation

Creating function tools dynamically.

Code

☎️ Outbound caller

Agent that makes outbound phone calls

Code

πŸ“‹ Structured output

Using structured output from LLM to guide TTS tone.

Code

πŸ”Œ MCP support

Use tools from MCP servers

Code

πŸ’¬ Text-only agent

Skip voice altogether and use the same code for text-only integrations

Code

πŸ“ Multi-user transcriber

Produce transcriptions from all users in the room

Code

πŸŽ₯ Video avatars

Add an AI avatar with Tavus, Hedra, Bithuman, LemonSlice, and more

Code

🍽️ Restaurant ordering and reservations

Full example of an agent that handles calls for a restaurant.

Code

πŸ‘οΈ Gemini Live vision

Full example (including iOS app) of Gemini Live agent that can see.

Code

Running your agent

Testing in terminal

python myagent.py console

Runs your agent in terminal mode, enabling local audio input and output for testing. This mode doesn't require external servers or dependencies and is useful for quickly validating behavior.

Developing with LiveKit clients

python myagent.py dev

Starts the agent server and enables hot reloading when files change. This mode allows each process to host multiple concurrent agents efficiently.

The agent connects to LiveKit Cloud or your self-hosted server. Set the following environment variables:

  • LIVEKIT_URL
  • LIVEKIT_API_KEY
  • LIVEKIT_API_SECRET

You can connect using any LiveKit client SDK or telephony integration. To get started quickly, try the Agents Playground.

Running for production

python myagent.py start

Runs the agent with production-ready optimizations.

Contributing

The Agents framework is under active development in a rapidly evolving field. We welcome and appreciate contributions of any kind, be it feedback, bugfixes, features, new plugins and tools, or better documentation. You can file issues under this repo, open a PR, or chat with us in LiveKit's Slack community.


LiveKit Ecosystem
LiveKit SDKsBrowser Β· iOS/macOS/visionOS Β· Android Β· Flutter Β· React Native Β· Rust Β· Node.js Β· Python Β· Unity Β· Unity (WebGL) Β· ESP32
Server APIsNode.js Β· Golang Β· Ruby Β· Java/Kotlin Β· Python Β· Rust Β· PHP (community) Β· .NET (community)
UI ComponentsReact Β· Android Compose Β· SwiftUI Β· Flutter
Agents FrameworksPython Β· Node.js Β· Playground
ServicesLiveKit server Β· Egress Β· Ingress Β· SIP
ResourcesDocs Β· Example apps Β· Cloud Β· Self-hosting Β· CLI

About

A powerful framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages