Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

VERBI - Voice Assistant 🎙️

GitHub StarsGitHub ForksGitHub IssuesGitHub Pull RequestsLicense

Motivation ✨✨✨

Welcome to the Voice Assistant project! 🎙️ Our goal is to create a modular voice assistant application that allows you to experiment with state-of-the-art (SOTA) models for various components. The modular structure provides flexibility, enabling you to pick and choose between different SOTA models for transcription, response generation, and text-to-speech (TTS). This approach facilitates easy testing and comparison of different models, making it an ideal platform for research and development in voice assistant technologies. Whether you're a developer, researcher, or enthusiast, this project is for you!

Features 🧰

  • Modular Design: Easily switch between different models for transcription, response generation, and TTS.
  • Support for Multiple APIs: Integrates with OpenAI, Groq, and Deepgram APIs, along with placeholders for local models.
  • Audio Recording and Playback: Record audio from the microphone and play generated speech.
  • Configuration Management: Centralized configuration in config.py for easy setup and management.

Project Structure 📂

voice_assistant/
├── voice_assistant/
│ ├── __init__.py
│ ├── audio.py
│ ├── api_key_manager.py
│ ├── config.py
│ ├── transcription.py
│ ├── response_generation.py
│ ├── text_to_speech.py
│ ├── utils.py
│ ├── local_tts_api.py
│ ├── local_tts_generation.py
├── .env
├── run_voice_assistant.py
├── setup.py
├── requirements.txt
└── README.md

Setup Instructions 📋

Prerequisites ✅

  • Python 3.10 or higher
  • Virtual environment (recommended)

Step-by-Step Instructions 🔢

  1. 📥 Clone the repository
 git clone https://github.com/PromtEngineer/Verbi.git
cd Verbi
  1. 🐍 Set up a virtual environment

Using venv:

 python -m venv venv
source venv/bin/activate # On Windows use `venv\Scripts\activate`

Using conda:

 conda create --name verbi python=3.10
conda activate verbi
  1. 📦 Install the required packages
 pip install -r requirements.txt
  1. 🛠️ Set up the environment variables

Create a .env file in the root directory and add your API keys:

 OPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
LOCAL_MODEL_PATH=path/to/local/model
  1. 🧩 Configure the models

Edit config.py to select the models you want to use:

 class Config:
# Model selection
TRANSCRIPTION_MODEL = 'groq'# Options: 'openai', 'groq', 'deepgram', 'fastwhisperapi' 'local'
RESPONSE_MODEL = 'groq'# Options: 'openai', 'groq', 'ollama', 'local'
TTS_MODEL = 'deepgram'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'# API keys and paths
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
GROQ_API_KEY = os.getenv("GROQ_API_KEY")
DEEPGRAM_API_KEY = os.getenv("DEEPGRAM_API_KEY")
LOCAL_MODEL_PATH = os.getenv("LOCAL_MODEL_PATH")

If you are running LLM locally via Ollama, make sure the Ollama server is runnig before starting verbi.

  1. 🔊 Configure ElevenLabs Jarvis' Voice
  • Voice samples here.
  • Follow this link to add the Jarvis voice to your ElevenLabs account.
  • Name the voice 'Paul J.' or, if you prefer a different name, ensure it matches the ELEVENLABS_VOICE_ID variable in the text_to_speech.py file.
  1. 🏃 Run the voice assistant
 python run_voice_assistant.py
  1. 🎤 Install FastWhisperAPI

    Optional step if you need a local transcription model

    Clone the repository

     cd..
    git clone https://github.com/3choff/FastWhisperAPI.git
    cd FastWhisperAPI

    Install the required packages:

     pip install -r requirements.txt

    Run the API

     fastapi run main.py

    Alternative Setup and Run Methods

    The API can also run directly on a Docker container or in Google Colab.

    Docker:

    Build a Docker container:

     docker build -t fastwhisperapi .

    Run the container

     docker run -p 8000:8000 fastwhisperapi

    Refer to the repository documentation for the Google Colab method: https://github.com/3choff/FastWhisperAPI/blob/main/README.md

  2. 🎤 Install Local TTS - MeloTTS

    Optional step if you need a local Text to Speech model

    Install MeloTTS from Github

    Use the following link to install MeloTTS for your operating system.

    Once the package is installed on your local virtual environment, you can start the api server using the following command.

     python voice_assistant/local_tts_api.py

    The local_tts_api.py file implements as fastapi server that will listen to incoming text and will generate audio using MeloTTS model. In order to use the local TTS model, you will need to update the config.py file by setting:

     TTS_MODEL = 'melotts'# Options: 'openai', 'deepgram', 'elevenlabs', 'local', 'melotts'

    You can run the main file to start using verbi with local models.

Model Options ⚙️

Transcription Models 🎤

  • OpenAI: Uses OpenAI's Whisper model.
  • Groq: Uses Groq's Whisper-large-v3 model.
  • Deepgram: Uses Deepgram's transcription model.
  • FastWhisperAPI: Uses FastWhisperAPI, a local transcription API powered by Faster Whisper.
  • Local: Placeholder for a local speech-to-text (STT) model.

Response Generation Models 💬

  • OpenAI: Uses OpenAI's GPT-4 model.
  • Groq: Uses Groq's LLaMA model.
  • Ollama: Uses any model served via Ollama.
  • Local: Placeholder for a local language model.

Text-to-Speech (TTS) Models 🔊

  • OpenAI: Uses OpenAI's TTS model with the 'fable' voice.
  • Deepgram: Uses Deepgram's TTS model with the 'aura-angus-en' voice.
  • ElevenLabs: Uses ElevenLabs' TTS model with the 'Paul J.' voice.
  • Local: Placeholder for a local TTS model.

Detailed Module Descriptions 📘

  • run_verbi.py: Main script to run the voice assistant.
  • voice_assistant/config.py: Manages configuration settings and API keys.
  • voice_assistant/api_key_manager.py: Handles retrieval of API keys based on configured models.
  • voice_assistant/audio.py: Functions for recording and playing audio.
  • voice_assistant/transcription.py: Manages audio transcription using various APIs.
  • voice_assistant/response_generation.py: Handles generating responses using various language models.
  • voice_assistant/text_to_speech.py: Manages converting text responses into speech.
  • voice_assistant/utils.py: Contains utility functions like deleting files.
  • voice_assistant/local_tts_api.py: Contains the api implementation to run the MeloTTS model.
  • voice_assistant/local_tts_generation.py: Contains the code to use the MeloTTS api to generated audio.
  • voice_assistant/__init__.py: Initializes the voice_assistant package.

Roadmap 🛤️🛤️🛤️

Here's what's next for the Voice Assistant project:

  1. Add Support for Streaming: Enable real-time streaming of audio input and output.
  2. Add Support for ElevenLabs and Enhanced Deepgram for TTS: Integrate additional TTS options for higher quality and variety.
  3. Add Filler Audios: Include background or filler audios while waiting for model responses to enhance user experience.
  4. Add Support for Local Models Across the Board: Expand support for local models in transcription, response generation, and TTS.

Contributing 🤝

We welcome contributions from the community! If you'd like to help improve this project, please follow these steps:

  1. Fork the repository.
  2. Create a new branch (git checkout -b feature-branch).
  3. Make your changes and commit them (git commit -m 'Add new feature').
  4. Push to the branch (git push origin feature-branch).
  5. Open a pull request detailing your changes.

Star History ✨✨✨

Star History Chart

About

A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, and Deepgram APIs, plus local models. Ideal for research and development in voice technology.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages