VinvAI
About
Vinv (Vibe Inverse) runs your services, finds issues, and verifies fixes - with zero code changes.
Explore
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
VinvAICommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"vinvai": {
"server": {
"command": "npx",
"args": [
"-y",
"vinv-mcp"
],
"env": {
"VINV_WORKSPACE": ""
}
}
}
}
}
McpServers
{
"server": {
"command": "npx",
"args": [
"-y",
"vinv-mcp"
],
"env": {
"VINV_WORKSPACE": ""
}
}
}
Transport
"stdio"
Package
"vinv-mcp"
Registry
"npm"
Vinv (Vibe Inverse) runs your services, finds issues, and verifies fixes — with zero code changes.
It connects runtime traces to the source code that produced them, hands that evidence to your agent, then runs the code again to verify the fix actually works. It even usesThompson samplingto decidehow muchruntime context to give the agent — because more is not always better.
Python first — services and APIs. TS & Go next.
No account. No API keys. No telemetry. Everything runs on your machine. Open source, Apache-2.0.
Vinv works three ways — as an editor extension, a CLI, or an MCP server for any agent. They share the same engines.
First run builds the engines — about 4 minutes: it compiles the Rust index and fetches a one-time ~500 MB local embedding model (uvandRustrequired). First trace lands about a minute after that; everything after is seconds.
pip install vinv # every engine as a console script # or run one with zero install: uvx --from vinv exerciser campaign <repo> --budget 20
Give Claude Code, Cursor, or any MCP client Vinv's tools. Install the engines above, then register the unified server —one global config; it finds your open workspace automatically via MCP roots:
Other clients: add{ "command": "npx", "args": ["-y", "vinv-mcp"] }undermcpServers.vinvin the client's MCP config. Seevinv-mcp— 16 tools: semantic search, dead code, fault localization, runtime values/slices/coverage, and the verify/optimize loop.
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines && cd ~/.vinv/engines && ./install.sh
git clone https://github.com/VinvAI/VinvAI $HOME\.vinv\engines; cd $HOME\.vinv\engines; .\install.ps1
Run · Test · Find — then Prove.Point Vinv at a Python repo; it does the rest — no code changes, no API keys.
- 🏃 Run— brings every service in your repo up under tracing with zero edits, capturing timings, arguments, return values and call trees from the real run.
- 🧪 Test— drives real requests through every endpoint (valid, boundary, negative, authenticated) and banks each response as a regression case.
- 🔎 Find— surfaces what actually broke or slowed down — server errors, crashes, latency hotspots and dead code — each tied to the exact source line.
- ✅ Prove— hands that evidence to the agent you already use (Claude Code, Cursor, Copilot…), then verifies its fix against acceptance tests writtenbeforethe fix that it never sees. A "faster" change that alters any output is auto-reverted.
Your agent is the only LLM — no new bill, no model picker, no provider keys. Everything runs on your machine.
84% of developers now use or plan to use AI coding tools. More of themactively distrustthe output (46%) than trust it (33%) — and distrust nearly doubled in a year (Stack Overflow 2025, 49k developers). You know why: the agent edits the wrong handler, invents return shapes, then grades its own homework while the server won't even start.
Or it gets stuck — test fails, agent edits the same function, test fails the same way, agent edits it again, burning your context window on "let me verify." Anthropic's own research documents agents "stuck in loops, repeating the same failed approach" when they lack codebase context.
Both failures have one root cause:the agent has never watched your code run.It argues from static text.
The industry automatedwritingand leftprovingentirely manual. Vinv automates the proving — and only then the finding and the fixing.
Receipts first — then how the loop produces them.
Vinv foundfour bugs and one performance probleminfastapi/full-stack-fastapi-template(~44k★). Same five issues, same prompts, Vinv grading every run:
One trial per condition — ademonstration, not a benchmark. Blind, the commodity model scored zero. Hand it the failing frame, the caller chain, and the real argument values, and it beats a stronger model guessing from static code.The evidence is what moved, not the weights.
On that same pristine template the optimization loop later detected — from live traces alone — that the app's default database pool makes requestsqueue for connection checkoutsunder concurrent load, dispatched the pool-sizing fix, and proved it: sustained-load median75.6ms → 41.2ms, 45.4% faster (95% CI [36.3%, 45.8%]), responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win wereauto-reverted— the accept landed only when the evidence did.
Same discipline, upstream on Hugging Face
That ishow come after it: oracles find the waste, your agent proposes the edit, paired-bootstrap + byte-identical replay decide accept or revert, and only then does anything go upstream. The rest of this README is the machinery behind those receipts.
TheTeststage above isn't one tester — it's a set of oracles, each hunting a different class of defect, all writing into the same findings and the same fix-dispatch path. You never need to think about them to use Vinv; open this if you want the full list and how the budget is spent.
The dispatcher is real, and it is a bandit.exerciser campaignallocatesone budgetacross everyarmedoracle by Thompson sampling over(target × technique × oracle)— rather than driving each one exhaustively. Cost ismeasured(wall-clock normalized to probe-equivalents plus subprocesses spawned), so an oracle that takes forty seconds to find what another finds in one loses. Credit is paidonce per defect signature, within a run and across runs, so a deterministic oracle can't re-earn credit for the same bug forever. Posteriors persist incampaign.json—which technique pays on your repois learned.
The campaign dispatchessixoracles (default_runners): crash/function, differential, fault, concurrency, HTTP, and environment — eacharmedonly when it applies to your repo (no--base-urland the HTTP oracle stays dark; no boundaries and the fault oracle does). Thedead-codeandruntime-analysisfamilies (leaks, hotspots, cache candidates) and thegolden-I/O baselinesin the table below run from the editor and the regression path, not this budget loop.
Unverified code runs behind acontainment ladder: a kernel-enforced OS sandbox (sandbox-exec/bwrap/unshare) where the host offers one, otherwise a process shim — always with a disposable repo copy, redirectedHOME/TMPDIR, blocked network and subprocess spawning. Tier is decided by aprobethat verifies a write outside the root really failed, never by a binary being onPATH. Postgres, Redis and S3 are substitutedinsidethe jail so code that needs them runs instead of failing to connect.
Static tools (Vulture,deadcode, Knip, ts-prune) can only prove"nothing statically references this."They can't see dynamic dispatch, feature flags, registries or environment drift — so they emit candidates a human has to adjudicate, which is why the cleanup never happens.
"No capture ever executed this. Here is what still references it, here is the traced neighbourhood it would wire back into, and here is what your agent thinks it is."
…
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



