View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
View jaedmunt's full-sized avatar
:shipit:
:shipit:

Block or report jaedmunt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaedmunt/README.md

Jaedon Munton

Building search engines, AI pipelines, and high-throughput systems

Metrics I care about: Cost per task · Precision · Data lineage · Recall

WebsiteEmailEmail DoublewordLinkedIn

Currently atDoubleword · PreviouslyChipHubEvvolve & Partners



About

I am a software engineer with a background in economics, interested in search and retrieval and the backend systems that support them.

  • Currently a Member of Technical Staff at Doubleword
  • Founder of Flux Search: freshness-first semantic search engine for developers and agents
  • Previously: investor-startup matching at Evvolve & Partners; agentic datasheet search at ChipHub (Nvidia Inception); credential platform at Certie (Oxford University Innovation); VC analyst scouting the MENA region at YAS Investments
  • Economics background with a focus on econometrics

Stack

LanguagesPythonGoTypeScriptSQL
Picking upRustZigC
Infra & BackendDockerKubernetesTerraformAWSAzureLinuxNginxPostgreSQLRedisFastAPISupabaseGitHub Actions
FrontendNext.jsTailwind CSSFigma
Data & MLPyTorchHugging FaceONNXJina AIpandasPolarsDuckDBClickHouseVespaWeaviateFAISSRay
ObservabilityGrafanaPrometheusLaminar
ToolsNeovimGitGitHub

Currently

Working on tagging, retrieval, and indexing decisions for Flux. I am particularly interested in the balance between throughput and latency, in prioritising cost per task, and in finding cost-effective methods to deliver better-quality and more explainable search systems.

  • Model pruning with PyTorch, run inside a flywheel loop of prune, eval, retrain, to remove parameters that contribute little to output quality.
  • Model quantization to shrink memory footprint and improve GPU utilisation at inference.
  • Fine-tuning small encoders, including multi-head (Hydra) architectures where one forward pass produces the typed spans that a decoder-phase model, an ensemble of larger models, or another heavier approach would otherwise require. I would describe myself as fairly hype-resistant, aiming for the most efficient (even if that means it's unsexy) option for the task even when that bucks the token-maxxing trend.
  • Matryoshka embeddings, quantization, and multi-modal embedding storage for high-quality retrieval across many ingested data types.
  • Evaluating tagger architectures, weighing LM-based approaches with a decoding step against encoder-only alternatives, to minimise the computational cost of delivering high-quality, explainable output.
  • Running reproducible, explainable eval suites across live-web data and static datasets, combining quantitative measures, qualitative review, LLM-as-judge grading, and commercial benchmarks.
  • Data curation, moderation, and crawl frontiers, covering filtering, deduping, and joining datasets so training and eval reflect real traffic, prioritising high-quality and authoritative sources, moderation that keeps unsafe content out of the index, and signed-URL frontiers that coordinate revisit scheduling and politeness across workers.
  • Data compaction and other methods to achieve token-efficient results.
  • Model serving through ONNX deployment paths, batching, SIMD in the hot inner loops, and hardware sized against realistic query patterns.
  • Choosing and designing the surrounding infrastructure for crawl, index, storage, cache, and observability, evaluating each swap on both engineering cost and unit economics.

Currently, I'm also exploring query expansion, query decomposition, and how to help LLMs know when they've arrived at a correct answer. In my free time I keep chipping away at Rust and reading retrieval papers. I also enjoy the operations side of team knowledge, stitching agents together to reduce friction between ideas and conversations, or setting up a note-taker that flags unclear moments and attributes decisions to the right people. Retrieval and context problems show up in many places, including search engines, team knowledge, agent workflows, and meeting notes. This way, I like to find where things break or could be improved and make efforts to improve them.

Where I'm opinionated

  • I prefer Rust and Go over Python on the hot path for tail latency, GIL-free concurrency, and more predictable memory behaviour.
  • I prefer measuring against realistic query patterns over headline benchmarks when evaluating tail latency and cost per task.
  • I prefer to make something work first, then make it work fast, then make it work for a lot of users.
  • For always-on systems like search, which balance precision and latency at high volume, the fundamentals matter early. Small decisions compound at scale.
  • Concretely, I do not want to degrade the search experience because a cost was miscalculated, or index material that is harmful or low quality.
  • As with music, if you can play an instrument slowly you can play it fast. Understanding a system carefully lets you scale it later.
  • The biggest dangers to a project are the ones that can end it entirely, like losing team integrity or failing to get to market on time. Team alignment matters more than any single technical decision.
  • I prefer building towards a big vision, however opaque, in small clear steps.

Interests

Together, these interests support the goal of building products where the technical choices add up to a fluid, commercially aligned user experience.

  • Backend systems, including balancing latency tradeoffs
  • Systems programming, mainly in Rust and Go
  • Econometrics, mostly DiD, network analysis, and measurement in tech
  • Finance, from venture to macro to equities

Big node, little node - Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.
XNV - Interactive XML navigator and filter with XPath-like queries. cargo install xnv / brew install xnv
Strike CLI - CLI tool for fast AI inference via Groq-hosted models, built as a formula and concept lookup.
Realms - Converts images into point clouds using a Facebook ML model. Built at Nvidia GTC.

Pinned Loading

  1. Tailored_SwiftTailored_SwiftPublic

    Want a high-accuracy voice clone quickly? Welcome to Tailored Swift! This collection offers phonetically balanced scripts covering the full range of sounds necessary for quality voice cloning. Desi…

    Python 35 4

  2. big-node-little-nodebig-node-little-nodePublic

    Distributed ML inference across a desktop RTX 3060 and a Raspberry Pi 4B, connected with Ray.

    Python

  3. xnvxnvPublic

    Interactive XML navigator and filter with XPath-like queries

    Rust

  4. 0xPlaygrounds/rig0xPlaygrounds/rigPublic

    ⚙️🦀 Build modular and scalable LLM Applications in Rust

    Rust 8.5k 950

  5. InftyAI/Awesome-LLMOpsInftyAI/Awesome-LLMOpsPublic

    🎉 An awesome & curated list of best LLMOps tools.

    Python 259 123

  6. doublewordai/control-layerdoublewordai/control-layerPublic

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 93 15