@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
@VinvAI

VinvAI

Tools for AI agents to test, fix and optimise your codebase
Vinv — tools for AI agents to test, fix and optimise your codebase.



Tools for AI agents to test, fix and optimise your codebase.

Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.

Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.


LicenseStarsVersionDownloadsTests100% local

vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai


// vision

Software that an AI writes should be judged by what it does, not by what it claims.

84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."

Both have one root cause: the agent has never watched your code run. It argues from static text.

So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.

The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.


// what we build

Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.

Runtime tracingZero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call.
Code Graph + semantic searchA Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers.
The context graphCode, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not.
Behavior exerciserDoesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case.
Verified fixesReplayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched.
Optimization with statisticsA speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted.
Agent babysittingA doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too.
MCP serversvinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI.
Vinv's nine stages around your coding agent: bring up, trace, index, map, exercise, find, dispatch, verify, learn
One command starts it. Vinv drives the other eight stages — every arrow is evidence, not a guess.

Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders

Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.


// does it work

Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:

SetupFixed
Cheap commodity model + Vinv context4 bugs + 1 optimization
Frontier model, working blind1 bug
Cheap commodity model, working blindnothing

A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.

The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.

On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.


// values

Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.

Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.

Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.

Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in ~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries, and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the agent CLI you configured, through your own auth.

Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.

Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.


// start here

VinvAI/VinvAIThe whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0.
InstallOne click, pick your editor — or code --install-extension VinvAI.VinvAI
Open VSX listingMarketplace page and reviews
vinv.aiThe short version, with the clips
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines &&cd~/.vinv/engines && ./install.sh

First run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.


// get involved

Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.

If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.


vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai

© 2026 VinvAI · Apache License 2.0 · Context beats model size.

Popular repositories Loading

  1. VinvAI VinvAIPublic

    Tools for AI agents to test, fix and optimise your codebase

    Python 46 5

  2. smolagents smolagentsPublic

    Forked from huggingface/smolagents

    🤗 smolagents: a barebones library for agents that think in code.

    Python

  3. .github .githubPublic

    Tools for AI agents to test, fix and optimise your codebase

  4. typer typerPublic

    Forked from fastapi/typer

    Typer, build great CLIs. Easy to code. Based on Python type hints.

    Python

  5. scikit-learn scikit-learnPublic

    Forked from scikit-learn/scikit-learn

    scikit-learn: machine learning in Python

    Python

  6. semantica semanticaPublic

    Forked from semantica-agi/semantica

    Graph-Native Infrastructure for Context and Accountable AI Systems

    Python

Repositories

Showing 10 of 15 repositories

Top languages

Loading…

Most used topics

Loading…