Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SnapIntel

Voice-to-Voice Assistant for Instant Insights from Screenshots

SnapIntel is a personal voice-to-voice assistant that provides immediate, actionable insights from the screenshots you decide to share. Whether you're solving an issue or looking for deeper understanding, SnapIntel is here to help.

This project is an open-source initiative that leverages Google Gemini to analyze images and provide responses. Various services are used to transcribe the user queries and generate spoken responses, including the local services FastWhisperAPI and FastXttsAPI.

If you find SnapIntel useful, please consider leaving a star ⭐ or donate.

Video Demo

Video Demo

Features

  • Easy and Intuitive Interface: Use voice-to-voice interactions for a seamless user experience.
  • Privacy-Focused Assistant: Maintain control over your data; decide what to share with a simple key combination press.
  • Instant Insights: Receive actionable information quickly from screenshots you choose to analyze.
  • Local Services Integration: Integrate with FastWhisperAPI and FastXttsAPI for localized query transcription and response vocalization.
  • Chat History: Records images and interactions within the session, enabling follow-up questions on images and recalling previous queries or responses.
  • Real-Time Session Logging: Automatically logs session history in a neatly formatted markdown file, accessible in real-time from the local logs folder.
  • Flexibility and Expandability: Built to adapt and grow with future enhancements and integrations.
  • Transcription Services: Support OpenAI, Groq, Deepgram, and FastWhisperAPI (Faster Whisper) for efficient transcription of user queries.
  • Speech Services: Support OpenAI, ElevenLabs, Cartesia, Deepgram, and FastXttsAPI (Coqui) for quick and natural-sounding vocalization of responses.

Requirements

  • Python 3.10 or greater
  • FFmpeg. Instructions on how to install it can be found here
  • FastWhisperAPI and FastXttsAPI offer local transcription and speech solutions. Their use is optional. For information on deployment and requirements of these services, please refer to their respective documentation.

Dependencies

This project depends on the following libraries:

  • pillow
  • python-dotenv
  • keyboard
  • requests
  • colorama
  • SpeechRecognition
  • google.generativeai
  • websocket-client
  • pyaudio
  • numpy

Installation

  1. Clone the repository:

    git clone https://github.com/3choff/SnapIntel.git
  2. Navigate to the project directory:

    cd SnapIntel
  3. Create a new environment:

    python3 -m venv SnapIntel
  4. Activate the virtual environment:

    • On Unix/Linux/macOS:

      source SnapIntel/bin/activate
    • On Windows:

      SnapIntel\Scripts\activate
  5. Install the required packages:

    pip install -r requirements.txt

Configuration

API keys

SnapIntel uses dotenv to set the API keys. Create a .env file in the root directory with your API keys. Follow the structure of the example.env file as a template.

Transcription and Speech services

The app supports multiple transcription and speech services right out of the box. You can select from the following options:

Transcription Services:

  • Deepgram
  • Openai
  • Groq
  • FastWhisperAPI, a local transcription API server using Faster Whisper.

Speech Services:

  • Deepgram
  • OpenAI
  • ElevenLabs
  • Cartesia (EXPERIMENTAL)
  • FastXttsAPI, a local speech API server using Coqui.

To change the transcription or speech service, simply edit the relevant variables in the Config.py file located in the services folder. The accepted choices are commented next to each variable.

In the same file, you can change other related variables such as voices and language.

Usage

To run the SnapIntel, use the following command:

python app.py

When the app starts, it will prompt you to either start a new session or resume a previous session stored in the history folder. After making your choice, you can interact with the LLM using these key combinations:

  • Press Ctrl+Alt+Space to capture and analyze the screen and invoke the voice assistant.
  • Press Ctrl+Space to ask a question without capturing a screenshot or to ask a follow-up question.
  • Press ESC to stop speech playback.
  • Press Ctrl+C to exit the script.

Support

If you find this project helpful and would like to support its development, there are several ways you can contribute:

  • Star: Consider leaving a star ⭐️ to increase the visibility of the project.
  • Support: Consider donate to support my work.
  • Contribute: If you're a developer, feel free to contribute to the project by submitting pull requests or opening issues.
  • Spread the Word: Share this project with others who might find it useful.

Your support means a lot and helps keep this project going. Thank you for your contribution!

Acknowledgements

This project is inspired by innovative features showcased by OpenAI in their demo of the upcoming features of ChatGPT, combining voice and vision capabilities to provide assistance and insights. The Verbi chatbot project and the Screen to Voice Tutorial of All About AI have significantly influenced this project, forming the foundation for its development. I recommend checking the links if you want to know more.

License

This project is licensed under the Apache License 2.0.

About

Voice-to-Voice Assistant for Instant Insights from Screenshots

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages