Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - redpluglab/tellama: Run private local LLMs and an Ollama/OpenAI-compatible API server on Android · GitHub
Skip to content

Repository files navigation

Tellama app icon

Tellama

Turn your Android phone into a private, Ollama-compatible AI server.
Run GGUF models on-device. Chat locally. Connect your PC, scripts, and agents through a small authenticated API.

Download the Android APK · 한국어 · API reference · Open SDK · Privacy

Android 10+Latest release v1.3.6Ollama compatibleOpenAI compatibleSDK license

Your phone. Your model. Your endpoint.

Tellama workspace dashboardTellama local AI chatTellama on-device model catalogTellama Ollama-compatible server

Workspace · Local chat · Model library · Private API server

Tellama is more than another chat screen. It turns hardware you already own into a reusable local AI endpoint:

  • Private by default: compatible GGUF models, chats, and long-term memory stay on your device.
  • Useful beyond the phone: trusted computers and agents can call Ollama- and OpenAI-compatible routes over your Wi-Fi.
  • Built for real devices: model guidance, measured memory and thermal status, resumable downloads, and safe model unloading protect limited mobile resources.
  • Lightweight research agent: summarize or compare public HTTPS pages with persistent sources, then choose whether the result belongs in long-term memory.
Mac / PC / Agent ── Ollama or OpenAI API ──▶ Android phone ──▶ Local GGUF model

The Android app is commercially licensed and its complete source remains private. The working Python and JavaScript clients, examples, and compatibility tests in sdk/ are open source under Apache-2.0.

Quick start

Upgrading from v1.2.0 on Android 10: install a signed v1.2.1-or-later APK manually once from the release page. Android retains your Tellama data and models. In-app updates initiated by v1.2.1 and later use the physically qualified update path.

  1. Install the latest APK.
  2. In Models, download and select a model that fits the phone.
  3. In Server, create an API key, choose Wi-Fi LAN, and start the server.
  4. Export the values shown by Tellama:
export TELLAMA_URL="http://PHONE_IP:11434"export TELLAMA_API_KEY="tlm_..."

List the exact model IDs installed on the phone:

curl "$TELLAMA_URL/api/tags" \
-H "Authorization: Bearer $TELLAMA_API_KEY"

Call the Ollama-compatible streaming chat route:

curl "$TELLAMA_URL/api/chat" \
-H "Authorization: Bearer $TELLAMA_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_API_TAGS", "messages": [{"role": "user", "content": "Hello from my Mac"}] }'

The response is newline-delimited JSON. OpenAI-style streaming is available at POST /v1/chat/completions using SSE. See the API compatibility reference for the exact supported routes and current limitations.

Open clients

No cloud account or package registry is required.

# Python 3.10+, standard library only
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
python3 sdk/python/example.py
# JavaScript, Node.js 18+
TELLAMA_URL="$TELLAMA_URL" TELLAMA_API_KEY="$TELLAMA_API_KEY" \
node sdk/javascript/example.mjs

The clients support model discovery, bounded generation options, Ollama NDJSON chat streaming, OpenAI SSE chat streaming, timeouts, and structured HTTP or stream errors. Run their contract suite with:

sdk/tests/run.sh

What the Android app includes

  • Phone-aware GGUF model catalog, import, selection, deletion, and per-model generation controls
  • Resumable large-model downloads with storage reserve, trusted mirrors, SHA-256 checks, and GGUF validation
  • Safe loaded-model deletion that stops active serving, unloads native memory, and then reclaims storage
  • Memory-aware model replacement that unloads the resident runtime before recalculating available RAM, reports selected/loading/ready states truthfully, and rolls back both selection and runtime after a failed load
  • Local streaming chat with timestamps, slash commands, voice input, and user-controlled long-term memory
  • User-initiated HTTPS page summaries and comparisons with persistent sources and explicit approval before saving a result to memory
  • Workspace dashboard for serving readiness, measured chat speed, RAM, storage, battery, and thermal guidance
  • Authenticated Wi-Fi LAN serving with one-time API keys, permission scopes, rate limiting, foreground status, and automatic stop on Wi-Fi loss
  • Non-blocking startup update reminders plus in-app download with SHA-256, package, version, and signing-certificate verification
  • Grouped commercial-grade settings with an app-wide, persistent 85%–130% text-size control
  • Hardened generation, model preflight, encrypted-data recovery, downloads, memory consent, and API reliability for friend testing
  • 15-language resources with TalkBack, large-text, landscape, tablet-width, and RTL layout support

Security boundary

  • LAN mode is off until the user starts it.
  • LAN requests require Authorization: Bearer <key>.
  • Full API keys are displayed once; Tellama stores their SHA-256 hashes.
  • Changing or losing Wi-Fi automatically stops the LAN server.
  • Do not expose port 11434 to the public internet or forward it from a router.
  • External API models have different privacy and billing conditions from Tellama's local server.
  • Web research makes direct HTTPS requests to the URLs you provide. Page text stays on the device when a local model is selected; with an external API model, that text is sent to the configured provider after an in-app warning.

Report security issues through SECURITY.md. Feature requests and SDK contributions are welcome; see CONTRIBUTING.md and the roadmap.

Licensing

  • Tellama Android application and brand assets: proprietary, all rights reserved.
  • Public SDK, examples, and compatibility tests under sdk/: Apache License 2.0.
  • Downloaded models: governed by each model publisher's license.

This repository intentionally does not contain the Tellama Android application source, native runtime implementation, signing material, or internal configuration.

About

Run private local LLMs and an Ollama/OpenAI-compatible API server on Android

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages