Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Onde Inference

Onde Inference

Run LLMs on-device with Onde Inference, with first-class support for Apple silicon.

crates.ioMaven CentralSwift Package Indexpub.devnpmWebsiteApp Store

Swift SDK · Kotlin Multiplatform SDK · Flutter SDK · React Native SDK · Website


In production

Onde is already shipping in real apps on the App Store and Google Play. Chat runs fully on-device, so there is no server round trip and no user data leaving the device. For SDK docs, platform notes, and setup details, see https://ondeinference.com/sdk. If you want to test downloads, model selection, or GGUF export before wiring the engine into app code, use Onde CLI.

Siti AI is the flagship open reference app — a private, on-device assistant built on Onde, open source under Apache-2.0. Its source is a complete, readable example of wiring the engine into a shipping Tauri app across macOS, iOS, and Android.

Download Siti AI on the App StoreGet Siti AI on Google Play


Model formats

Onde can load GGUF models and UQFF models through the same chat engine. UQFF is mistral.rs' native pre-quantized format. Point model_id at the UQFF export — the repository or local directory holding the shards, residual.safetensors, config.json, and the tokenizer — and name the first shard (or a shorthand such as q4k) in files. Use the UQFF repository, not the original unquantized one; the export is self-contained and everything is resolved relative to it.

use onde::inference::{ChatEngine,UqffModelConfig};let engine = ChatEngine::new();
engine
.load_uqff_model(UqffModelConfig{model_id:"mistralrs-community/gemma-4-E4B-it-UQFF".into(),files:vec!["q4k-0.uqff".into()],display_name:"Gemma 4 E4B (UQFF Q4K)".into(),approx_memory:"~2.5 GB (UQFF Q4K)".into(),chat_template:None,},None,None,).await?;

For sharded UQFFs, passing the first shard is enough; mistral.rs discovers sibling shards with the same prefix.


License

Onde is dual-licensed under MIT and Apache 2.0. You may use it under either license at your option.

Dependency attribution

DependencyLicenseAuthor
mistral.rsMITEric Buehler
UniFFIMPL-2.0Mozilla
tokioMITTokio contributors

Model licenses

Models downloaded by Onde have their own licenses independent of this crate. By using Onde, you are also subject to the license of the model you load:

ModelSizeLicenseCommercial use
Qwen 2.5 1.5B Instruct (GGUF Q4_K_M)~941 MBQwen Community License✅ with conditions¹
Qwen 2.5 3B Instruct (GGUF Q4_K_M)~1.93 GBQwen Community License✅ with conditions¹
Qwen 2.5 Coder 7B Instruct (GGUF Q4_K_M)~4.4 GBQwen Community License✅ with conditions¹
Qwen 3 1.7B (GGUF Q4_K_M)~1.3 GBApache 2.0
Qwen 3 4B (GGUF Q4_K_M)~2.7 GBApache 2.0
Qwen 3 8B (GGUF Q4_K_M)~5 GBApache 2.0
Qwen 3 14B (GGUF Q4_K_M)~8.4 GBApache 2.0
DeepSeek Coder 6.7B Instruct (GGUF Q4_K_M)~3.8 GBDeepSeek License v1.0✅ with conditions²

¹ Qwen Community License conditions: no training of competing models, attribution required, no misrepresentation of origin. Organisations with more than 100 million monthly active users must obtain a separate commercial licence from Alibaba Cloud.

² DeepSeek License v1.0 conditions: use-based restrictions apply (see Attachment A of the license). Prohibits military use, generation of disinformation, and certain other uses. Governing law is PRC law.

Onde's own license (MIT OR Apache-2.0) is independent of these model licenses. If you build an application on top of Onde, you are responsible for complying with the license of whichever model your users load.


Copyright

© 2026 Splitfire AB (Onde Inference).

About

Open-source Llm for Apple silicon devices.

Topics

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages