Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+License: MITPowered by OllamaUI: WebView2 / PyQt6Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher"Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
WW["Wake word"] --> ROUTE
HK["Ctrl+Space"] --> ROUTE
TX["Typed input"] -.skips recording.-> STT
ROUTE(( )) --> REC["Recorder: VAD"]
REC --> STT["STT: faster-whisper\n(live partial + final)"]
STT --> INSTANT{"Predefined\nphrase?"}
INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
CMD -- yes --> SYS["OS action\n(no shell, ever)"]
CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
SYS --> TTS
LLM --> TTS
TTS --> IDLE["back to idle"]
style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

StageWhat runs
Wake wordopenwakeword (ONNX), continuous, low CPU
RecordingRMS-based voice activity detection
Transcriptionfaster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responsesDictionary lookup (commands/instant.py): never touches the LLM
Instant commandsRegex router -> os.startfile / ctypes / list-form subprocess (never a shell)
ConversationLocal Ollama, or Claude if configured
Speech outedge-tts (online voices) or pyttsx3 (fully offline)
OverlayAlways-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

MethodResult
Say "Hey Sam" (default wake word)Orb lights up, starts listening
Press Ctrl+SpaceSame, no wake word needed
Press Ctrl+Shift+SpaceOpens a text box under the orb instead
Click the orbSame as the text hotkey
Ctrl + drag the orbMoves it: position is remembered
Right-click the tray iconSettings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design."shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
trigger: ctrl+space # hold to speaktext_input: ctrl+shift+space # open the typed-input boxui:
orb:
layer: auto # auto (bottom until called, then on top) | topmost | normalclick_through: trueidle_animation: trueidle_fps: 12# SAM runs 24/7 - low idle costactive_fps: 60wake_word:
model: assets/models/hey_sam.onnxthreshold: 0.5# lower = triggers more easilystt:
model: small # tiny | base | small | medium | large-v3device: cpu # or cuda, with a working CUDA + cuDNN setupllm:
ollama:
model: qwen2.5:3bautostart: true # SAM starts "ollama serve" itself if it isn't runningstop_on_exit: false # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()patterns.append((
re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
lambdam: system.get_cpu_temperature()
))
# commands/system.pydefget_cpu_temperature() ->str:
"""Handlers never raise - the router already wraps calls in try/except."""try:
# Real OS side effects use list-form subprocess or os.startfile (never shell=True)
...
returnf"Your CPU is at {celsius:.1f} degrees."exceptException:
return"Sorry, I couldn't read the CPU temperature."

Troubleshooting

SymptomLikely causeFix
No LLM engine found in the logOllama is not installed, or the model is not pulledCheck the log for details; run ollama pull qwen2.5:3b
Wake word does not triggerThreshold too high, or wrong micLower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silenceWhisper hallucination behavior on background noiseRaise audio.silence_threshold
Ctrl+Space does nothingHotkey hook needs elevated access on some windowsRun SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeysTwo SAM processes runningSAM allows one instance only (named mutex) - check the tray before starting another
Settings will not saveInstalled build config lives in %APPDATA%\SAM\config.yamlEdit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/ icon, activation chime, wake word model
├── audio/ wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/ regex router + OS side effects + clipboard reader
├── core/
│ ├── app.py AppController: state machine, wires everything together
│ ├── config.py DEFAULTS + config.yaml loader/saver
│ ├── paths.py dev vs. frozen-exe path resolution, single-instance lock
│ ├── code_parser.py extracts code blocks from LLM replies to the Desktop
│ └── installer_steps.py SAM.exe --install-models (used by installer)
├── llm/ LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│ ├── web/ Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│ ├── web_settings.py pywebview settings host with native Win32 single instance
│ ├── orb.py · caption.py The always-on overlay
│ ├── toast.py Ephemeral HUD toast for zero-LLM feedback
│ ├── win32.py click-through, z-order, foreground-focus helpers
│ └── tray.py System tray integration
├── installer/SAM.iss Inno Setup script
├── tools/make_icon.py regenerates assets/icon.ico from the orb design
├── SAM.spec PyInstaller build spec
├── config.example.yaml committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text.faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step.shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo.config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages