Semantic Sift
About
# 🔍 Semantic-Sift **The Reasoning-First Middleware for High-Fidelity Agentic Workflows.** [](https://github.com/luismichio/semantic-sift/actions/workflows/ci.yml)…
Explore
| Feature | Python MCP Server | Rust Sift-Core (Sidecar) |
| :--- | :---: | :---: |
| Heuristic Log Sifting | ✅ | ✅ (Native) |
| Semantic Compression | ✅ (PyTorch) | ✅ (ONNX) |
| Multi-Modal Ingestion | ✅ (via [multi-modal]) | ❌ (Text Only) |
| Supported Formats | .pdf, .xlsx, .docx, .html, .txt | .txt, .log, .out (Text) |
| Startup Latency | 3-5 seconds | ~10ms |
| Binary Size | ~1.5GB (with models) | ~15MB |
> Note: For native apps like Meechi, we recommend a Tiered Ingestion strategy: use the app's frontend (e.g., pdf.js) to extract text, then pipe it to the Rust sidecar for high-speed semantic sifting.
Usage:
```bash
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
Semantic SiftCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
pip install "semantic-sift[neural]"
Option A: Quick Install (PyPI)
> ℹ️ What you get: The PyPI wheel includes the pre-compiled sift-core Rust binary — no Rust toolchain required. The [neural] extra adds PyTorch (~1.5 GB) for large-payload fallback using LLMLingua-2; [multi-modal] adds MarkItDown for PDF/DOCX/XLSX ingestion. Expect several minutes for the first install due to PyTorch download size.
bashuv venv
Choosing the right Python path for your MCP configuration is critical for stability:
| Setup Type | Path Example | Pros | Cons |
| :--- | :--- | :--- | :--- |
| Dedicated Venv (Win) | .../semantic-sift/venv312/Scripts/python.exe | Isolated dependencies, no torch version conflicts. | Slightly more disk space. |
| Dedicated Venv (Mac/Linux) | .../semantic-sift/venv312/bin/python | Same isolation benefit on Unix. | Same. |
| Global Python | C:/Users/User/AppData/Local/.../python.exe | Shared libraries, fast setup. | High risk of version conflicts (e.g., transformers mismatches). |
Recommendation: Always use the Dedicated Venv path in your mcp_config.json to ensure the sifting kernel is isolated and reliable.
> Note on Orchestration: Semantic-Sift is an "Intelligence Kernel." For complex multi-tool workflows, we strongly recommend installing Context-Pipe, the universal switchboard that natively routes data to Semantic-Sift without blocking your IDE.
For development tools (mypy, pytest):
uv pip install -e .[dev]
> Rust binary for editable installs: pip install -e . skips the Rust compile step, so sift-core won't be on your PATH. Instead of compiling from source, download the pre-built binary for your platform from the matching GitHub release in one command:
>
> python scripts/fetch_sift_core.py
> > This places
sift-core[.exe] directly into your active environment's Scripts/bin directory. Re-run it whenever you bump the version.Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"semantic sift": {
"semantic-sift": {
"command": "semantic-sift",
"args": [],
"env": {
"SIFT_ALLOW_GLOBAL_READS": "false"
}
}
}
}
}
McpServers
{
"semantic-sift": {
"command": "semantic-sift",
"args": [],
"env": {
"SIFT_ALLOW_GLOBAL_READS": "false"
}
}
}
The Reasoning-First Middleware for High-Fidelity Agentic Workflows.
> "It saves tokens while preserving context - maximizing reasoning, minimizing hallucination."
Semantic-Sift is a local Model Context Protocol (MCP) server that acts as an intelligent "Sanitation Tier" between your raw data and your AI’s context window.
While modern LLMs have massive context windows, their reasoning accuracy often degrades as noise increases. Semantic-Sift solves this by distilling technical logs, long-form documents, and chat histories into high-density context using LLMLingua-2. It treats your context window as a precious resource—optimizing for Signal-to-Noise Ratio (SNR) so your models spend more time reasoning and less time navigating boilerplate.
🧠 Philosophy: The Studio of Two
Semantic-Sift is grounded in the Studio of Two philosophy: the belief that the future of engineering is a high-fidelity partnership between a human architect and a sovereign AI sidecar. By managing the friction of raw data ingestion, Sift allows this "Studio" to focus on building systems, not just applying patches. It acts as a cognitive filter that ensures both you and your agent are collaborating on the cleanest, most relevant representation of the technical truth.---
⚡ Quickstart (60 seconds)
```bash
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



