Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

1,678 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ONNX Runtime GenAI

Status

Latest version

Nightly Build

Description

Run generative AI models with ONNX Runtime. This API gives you an easy, flexible and performant way of running LLMs on device. It implements the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling.

ONNX Runtime GenAI powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit.

See documentation at the ONNX Runtime website for more details.

Support matrixSupported nowUnder developmentOn the roadmap
Model architecturesAMD OLMo
ChatGLM
DeepSeek
ERNIE 4.5
Fara
Gemma
gpt-oss
Granite
Granite MoE Hybrid
HunYuan Dense V1
InternLM2
Llama
Mistral
Nemotron
Phi (language + vision)
Qwen (language + vision)
SmolLM3
Whisper
Stable diffusionMulti-modal models
APIPython
C#
C/C++
Java ^
Objective-C
O/SLinux
Windows
Mac
Android
iOS
Architecturex86
x64
arm64
Hardware AccelerationCPU
CUDA
DirectML
NvTensorRtRtx (TRT-RTX)
OpenVINO
QNN
WebGPU
AMD GPU
FeaturesMulti-LoRA
Continuous decoding
Constrained decoding
Speculative decoding

^ Requires build from source

Installation

See installation instructions or build from source

Sample code for Phi-3 in Python

  1. Download the model

    huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
  2. Install the API

    pip install numpy
    pip install --pre onnxruntime-genai
  3. Run the model

    importonnxruntime_genaiasogmodel=og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
    tokenizer=og.Tokenizer(model)
    stream=tokenizer.create_stream()
    # Set the max length to something sensible by default,# since otherwise it will be set to the entire context lengthsearch_options= {}
    search_options['max_length'] =2048search_options['batch_size'] =1chat_template='<|user|>\n{input} <|end|>\n<|assistant|>'text=input("Input: ")
    ifnottext:
    print("Error, input cannot be empty")
    exit()
    prompt=f'{chat_template.format(input=text)}'input_tokens=tokenizer.encode(prompt)
    params=og.GeneratorParams(model)
    params.set_search_options(**search_options)
    generator=og.Generator(model, params)
    print("Output: ", end='', flush=True)
    try:
    generator.append_tokens(input_tokens)
    whilenotgenerator.is_done():
    generator.generate_next_token()
    new_token=generator.get_next_tokens()[0]
    print(stream.decode(new_token), end='', flush=True)
    exceptKeyboardInterrupt:
    print(" --control+c pressed, aborting generation--")
    print()
    delgenerator

Choose the correct version of the examples

Due to the evolving nature of this project and ongoing feature additions, examples in the main branch may not always align with the latest stable release. This section outlines how to ensure compatibility between the examples and the corresponding version.

Stable version

Install the package according to the installation instructions. For example, install the Python package.

pip install onnxruntime-genai

Get the version of the package

Linux/Mac:

pip list | grep onnxruntime-genai

Windows:

pip list | findstr "onnxruntime-genai"

Then, check out the version of the examples that corresponds to that release.

# Clone the repo
git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai
# Checkout the branch for the version you are using
git checkout v0.11.5
cd examples

Nightly version (main branch)

Checkout the main branch of the repo

git clone https://github.com/microsoft/onnxruntime-genai.git &&cd onnxruntime-genai

Build from source, using these instructions. For example, to build the Python wheel:

python build.py

Navigate to the examples folder in the main branch.

cd examples

To install the nightly Python build:

# Change onnxruntime-genai to the Python package you want to install
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai

Roadmap

See the Discussions to request new features and up-vote existing requests.

Data/Telemetry

This project may collect usage data and send it to Microsoft to help improve our products and services. See the privacy statement for details.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.opensource.microsoft.com.

See DEVELOPMENT.md for how to build, test, and lint the library from source.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Linting

This project enables lintrunner for linting. You can install the dependencies and initialize with

pip install -r requirements-lintrunner.txt
lintrunner init

This will install lintrunner on your system and download all the necessary dependencies to run linters locally.

To format local changes:

lintrunner -a

To format all files:

lintrunner -a --all-files

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

About

Generative AI extensions for onnxruntime

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages