knowledge-rag
About
Local RAG system for Claude Code with hybrid search (semantic + BM25), cross-encoder reranking, markdown-aware chunking, 9 file formats, file watcher, and 12 MCP tools. Zero external servers. pip install knowledge-rag
Explore
Every RAG framework claims "production-ready." Here is what knowledge-rag shipsin the OSS core, verified by regression tests, that competitors either paywall, plugin-ify, or simply don't have.
- Bearer token authon SSE / HTTP transports — constant-time comparison (hmac.compare_digest), RFC 6750 challenge, 401 fenced withWWW-Authenticateheader
- Path traversal + symlink escape defenses—validate_path_withinguarding 6 CRUD tools (CWE-22, CWE-59)
- Prompt injection 3-layer defense— sentinel neutralization + provenance fence +external_sourceflag (OWASP LLM01:2025)
- OpenSSF Best Practices badgeverified ·CodeQLweekly scan ·Bandit + Semgrep + Gitleaks + pip-auditon every PR
- PyPI Trusted Publishingvia OIDC (zero long-lived tokens in CI)
- Prometheus/metricsendpoint— custom histogram buckets tuned for RAG (p95 ≤ 10ms fast-path targets), 7 canonical metrics via@instrumentdecorator on all 13 tools
- Rate limiting— thread-safe sliding-window counter, per-client RPM + burst, zero overhead when disabled
- Health probes—GET /healthand/healthzreturning{status, version, uptime_seconds, cache}in front of the auth middleware (probes always succeed)
- Structured JSON logging— opt-in viaserver.logging.format: "json", one JSON object per record ready for ELK / Loki / Datadog / CloudWatch
- Public benchmark dashboardon GitHub Pages
- SSE / streamable-http transport— 1 server serves N MCP clients, ChromaDB WAL mode enabled automatically, shared embedding model + query cache
- BM25 inverted-index—128× fasterthan linear scan (custom implementation, replacesrank-bm25)
- FTS5 SQLite fast-path(opt-in, ADR-002/003/006/008) — <10ms cold, <2ms hot on lexical queries
- Cross-encoder reranking— Xenova/ms-marco-MiniLM-L-6-v2, +1.88pp Recall@10 (p<0.001)
- GPU CUDA 12with auto DLL discovery + graceful CPU fallback
- Query cache— LRU + 5-min TTL, cuts p95 latency ~40%
- Zero-downtime reindex— staging populate + validation + atomic swap + durable metadata rollback
- Async background reindexwithget_reindex_status()polling
- Nightly chaos injection— HuggingFace Hub offline · ONNX zero-byte replay · watchdog crash recovery (3 scenarios intests/chaos/)
- 50 000-iteration soak test— proves no memory leak after 1h of continuous queries (KNOWLEDGE_RAG_SOAK_ITERATIONS=50000)
- Mutation testing(mutmut) oninstance_lock+preflight— catches tests that are too weak
- Determinism check— full test suite × 3, catches flakes
- Backwards-compat frozen— 13 MCP tool parameter names guarded bytests/test_backwards_compat.py+ legacy YAML fixtures (v3.6.0 / v3.7.0) still parse
- API surface AST diff—check_api_surface.pyblocks any breaking change at PR time
- 9-cell CI matrix— Linux + Windows + macOS × 3.11 + 3.12 + 3.13
Preset:](https://github.com/lyonzin/knowledge-rag/blob/HEAD/docs/API.md)[cybersecurity.yaml· 8 categories · 200+ routing keywords · 69 query expansions
Ingest MITRE ATT&CK, threat reports, exploit writeups, incident reports. Search from Claude Code withsearch_knowledge("privilege escalation windows")and get instant recall across your entire corpus. Air-gapped — nothing leaves the laptop.
Pick your integration path — knowledge-rag ships the same server through every channel.
Every RAG framework claims "production-ready." Here is what knowledge-rag shipsin the OSS core, verified by regression tests, that competitors either paywall, plugin-ify, or simply don't have.
- Bearer token authon SSE / HTTP transports — constant-time comparison (hmac.compare_digest), RFC 6750 challenge, 401 fenced withWWW-Authenticateheader
- Path traversal + symlink escape defenses—validate_path_withinguarding 6 CRUD tools (CWE-22, CWE-59)
- Prompt injection 3-layer defense— sentinel neutralization + provenance fence +external_sourceflag (OWASP LLM01:2025)
- OpenSSF Best Practices badgeverified ·CodeQLweekly scan ·Bandit + Semgrep + Gitleaks + pip-auditon every PR
- PyPI Trusted Publishingvia OIDC (zero long-lived tokens in CI)
- Prometheus/metricsendpoint— custom histogram buckets tuned for RAG (p95 ≤ 10ms fast-path targets), 7 canonical metrics via@instrumentdecorator on all 13 tools
- Rate limiting— thread-safe sliding-window counter, per-client RPM + burst, zero overhead when disabled
- Health probes—GET /healthand/healthzreturning{status, version, uptime_seconds, cache}in front of the auth middleware (probes always succeed)
- Structured JSON logging— opt-in viaserver.logging.format: "json", one JSON object per record ready for ELK / Loki / Datadog / CloudWatch
- Public benchmark dashboardon GitHub Pages
- SSE / streamable-http transport— 1 server serves N MCP clients, ChromaDB WAL mode enabled automatically, shared embedding model + query cache
- BM25 inverted-index—128× fasterthan linear scan (custom implementation, replacesrank-bm25)
- FTS5 SQLite fast-path(opt-in, ADR-002/003/006/008) — <10ms cold, <2ms hot on lexical queries
- Cross-encoder reranking— Xenova/ms-marco-MiniLM-L-6-v2, +1.88pp Recall@10 (p<0.001)
- GPU CUDA 12with auto DLL discovery + graceful CPU fallback
- Query cache— LRU + 5-min TTL, cuts p95 latency ~40%
- Zero-downtime reindex— staging populate + validation + atomic swap + durable metadata rollback
- Async background reindexwithget_reindex_status()polling
- Nightly chaos injection— HuggingFace Hub offline · ONNX zero-byte replay · watchdog crash recovery (3 scenarios intests/chaos/)
- 50 000-iteration soak test— proves no memory leak after 1h of continuous queries (KNOWLEDGE_RAG_SOAK_ITERATIONS=50000)
- Mutation testing(mutmut) oninstance_lock+preflight— catches tests that are too weak
- Determinism check— full test suite × 3, catches flakes
- Backwards-compat frozen— 13 MCP tool parameter names guarded bytests/test_backwards_compat.py+ legacy YAML fixtures (v3.6.0 / v3.7.0) still parse
- API surface AST diff—check_api_surface.pyblocks any breaking change at PR time
- 9-cell CI matrix— Linux + Windows + macOS × 3.11 + 3.12 + 3.13
Preset:](https://github.com/lyonzin/knowledge-rag/blob/HEAD/docs/API.md)[cybersecurity.yaml· 8 categories · 200+ routing keywords · 69 query expansions
Ingest MITRE ATT&CK, threat reports, exploit writeups, incident reports. Search from Claude Code withsearch_knowledge("privilege escalation windows")and get instant recall across your entire corpus. Air-gapped — nothing leaves the laptop.
📊 How knowledge-rag compares to other RAG frameworks
We audited16 popular RAG frameworks and platforms(LlamaIndex, LangChain, ChromaDB, Weaviate, Qdrant, RAGFlow, LightRAG, DSPy, GraphRAG, Haystack, RAG-Anything, kotaemon, txtai, llmware, Dify, open-webui, FastGPT) so you can pick honestly.
Legend:✅ built-in · 🟡 plugin / paid tier / partial · ❌ not available · ⚠️ license or default concern
The 5 dimensions where knowledge-rag is unique:health probes + JSON logging + Prometheus + rate limit + bearer authsimultaneously built-in on an OSS RAG-focused MCP server. Zero-downtime reindex + async background reindex + nightly chaos/soak/mutation are documented on nobody else's README.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.


