Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Doubleword Inference Stack

A Helm chart for deploying LLMs. This stack is a lightweight, transparent framework to allow you to deploy any model on any inference engine with minimal configuration. It is designed to be flexible, allowing you to easily switch between different models and inference engines without changing your client code.

The spirit of this project is to produce a framework that solves distributed serving of LLMs in a non-specific way to the inference engine running the weights.

We achieve this by deploying the Onwards AI Gateway with configurable LLM model groups. This chart provides a transparent, unified interface for routing requests to multiple inference engines like vLLM, SGLang, TensorRT-LLM, and others. It also allows you to set human readable model names that map to backend services, making it easy to switch between models without changing client code.

If you want to go beyond what's available here for high-throughput deployments, contact us at hello@doubleword.ai.

Architecture Overview

Architecture Diagram

The stack consists of:

  • Onwards Gateway: The API gateway that routes requests to different inference engines.
  • Model Groups: A grouping of kubernetes resources that represent a deployment of an inference engine. Each model group can only have one active model at a time, but can have custom numbers of replicas.
  • Inference Engines: Backend services like vLLM, SGLang, TensorRT-LLM that handle the actual inference requests.

Getting Started

Install from OCI Registry

# Add the Helm repository (OCI format)
helm pull oci://ghcr.io/doublewordai/inference-stack
# Install with default configuration
helm install my-inference-stack oci://ghcr.io/doublewordai/inference-stack

Install from Source

# Clone the repository
git clone https://github.com/doublewordai/inference-stack.git
cd inference-stack
# Install with default values
helm install my-inference-stack .# Or install with custom values
helm install my-inference-stack . -f my-values.yaml

Basic Configuration

The most important configuration is defining your model groups. Each model group represents a deployment of an inference engine:

modelGroups:
# vLLM deployment serving Llama modelsvllm-llama:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "llama"
- "llama-3.1-8b-instruct"modelName: "meta-llama/Meta-Llama-3.1-8B-Instruct"command:
- "vllm"
- "serve"
- "--model"
- "meta-llama/Meta-Llama-3.1-8B-Instruct"# SGLang deployment serving Qwen models sglang-qwen:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "qwen"
- "qwen-2.5-7b-instruct"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"

GPU Configuration

For GPU-enabled inference:

modelGroups:
vllm-llama:
enabled: trueresources:
limits:
nvidia.com/gpu: 2memory: 16Girequests:
nvidia.com/gpu: 1memory: 8GinodeSelector:
accelerator: nvidia-tesla-v100tolerations:
- key: nvidia.com/gpuoperator: Existseffect: NoSchedule

Persistent Storage for Model Caching

modelGroups:
vllm-llama:
# Single persistent volumepersistentVolumes:
- name: model-cachesize: 100GistorageClass: fast-ssdmountPath: /root/.cache/huggingfaceaccessModes:
- ReadWriteOnce# Multiple persistent volumes (advanced)# persistentVolumes:# - name: model-cache# size: 50Gi# mountPath: /root/.cache/huggingface# - name: model-weights# size: 200Gi# mountPath: /models# storageClass: fast-ssd

Deployment Strategy Configuration

modelGroups:
vllm-llama:
# Rolling update strategy for zero-downtime deploymentsstrategy:
type: RollingUpdaterollingUpdate:
maxSurge: 1# Allow 1 extra pod during updatesmaxUnavailable: 0# Never take pods down (blue-green)# Scale replicas for high availabilityreplicaCount: 3

External Access via Ingress

ingress:
enabled: trueclassName: nginxannotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"hosts:
- host: inference.example.compaths:
- path: /pathType: Prefixtls:
- secretName: inference-tlshosts:
- inference.example.com

Example Configurations

This framework can be used to deploy any set inference engines, we offer these as examples to get you started quickly. The examples/ directory contains ready-to-use configurations:

  • single-vllm.yaml - Single vLLM model deployment with persistent caching
  • single-sglang.yaml - SGLang deployment optimized for structured generation
  • single-tensorrt-llm.yaml - TensorRT-LLM with model compilation init container
  • multi-engine.yaml - Complete multi-engine setup (vLLM + SGLang + TensorRT-LLM)

Use any example as a starting point, for example:

helm install my-stack . -f examples/single-vllm.yaml

Usage Examples

Deploying Multiple Models

# values.yamlmodelGroups:
# Code generation with vLLMvllm-codegen:
enabled: trueimage: vllm/vllm-openaitag: latestmodelAlias:
- "codegen"
- "code-generation"modelName: "Salesforce/codegen-2B-multi"command:
- "vllm"
- "serve"
- "--model"
- "Salesforce/codegen-2B-multi"# Chat models with SGLangsglang-chat:
enabled: trueimage: lmsysorg/sglangtag: latestmodelAlias:
- "chat"
- "qwen-chat"modelName: "Qwen/Qwen2.5-7B-Instruct"command:
- "python"
- "-m"
- "sglang.launch_server"
- "--model-path"
- "Qwen/Qwen2.5-7B-Instruct"# High-performance with TensorRT-LLMtensorrt-llm:
enabled: trueimage: nvcr.io/nvidia/tritonservertag: 24.01-trtllm-python-py3modelAlias:
- "optimized"
- "fast-inference"modelName: "mistralai/Mistral-7B-Instruct-v0.2"command:
- "tritonserver"
- "--model-repository=/models"

Using the API

Once deployed, you can use the standard OpenAI API format:

# Get available models
curl -X GET http://inference.example.com/v1/models
# Generate completion
curl -X POST http://inference.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama-3.1-8b-chat", "messages": [{"role": "user", "content": "Hello!"}] }'

Development

Prerequisites

  • Helm 3.13+
  • Kubernetes 1.20+
  • (Optional) GPU nodes with NVIDIA device plugin

Running Tests

# Install test dependencies
helm plugin install https://github.com/helm-unittest/helm-unittest.git
# Run linting
helm lint .# Run unit tests 
helm unittest .# Test template rendering
helm template test-release .

Local Development

# Test with different configurations
helm template test-release . \
--set modelGroups.vllm-llama.enabled=true \
--set ingress.enabled=true \
--debug
# Install locally
helm install test-release . \
--set modelGroups.vllm-llama.enabled=true \
--dry-run

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Reach out at: hello@doubleword.ai!

About

The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages