Sourcerer

by st3v3nmw

303 downloads Not rated yet
GitHub

About

MCP for semantic code search & navigation that reduces token waste

Explore

- Semantic search by concept and functionality
- Retrieve specific code chunks by stable ID
- Uses Tree-sitter for AST-based code parsing
- Automatically re-indexes changed files via file watching
- Respects .gitignore rules
- Stores embeddings persistently in .sourcerer/db/

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Sourcerer
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

- OpenAI API Key: Required for generating embeddings (local embedding support planned)
- Git: Must be a git repository (respects .gitignore files)
- Add .sourcerer/ to .gitignore: This directory stores the embedded vector database

semantic_search

Find code by concept/functionality

get_source_code

Retrieve specific chunks by ID

index_workspace

Manually trigger re-indexing

get_index_status

Check indexing progress

- semantic_search: Find code by concept/functionality
- get_source_code: Retrieve specific chunks by ID
- index_workspace: Manually trigger re-indexing
- get_index_status: Check indexing progress

This approach allows AI agents to find relevant code without reading entire files,
dramatically reducing token usage and cognitive load.

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "sourcerer": {
            "sourcerer": {
                "command": "sourcerer",
                "env": {
                    "OPENAI_API_KEY": "your-openai-api-key",
                    "SOURCERER_WORKSPACE_ROOT": "/path/to/your/project"
                }
            }
        }
    }
}

McpServers

{
    "sourcerer": {
        "command": "sourcerer",
        "env": {
            "OPENAI_API_KEY": "your-openai-api-key",
            "SOURCERER_WORKSPACE_ROOT": "/path/to/your/project"
        }
    }
}
An MCP server for semantic code search & navigation that helps AI agents work efficiently without burning through costly tokens. Instead of reading entire files, agents can search conceptually and jump directly to the specific functions, classes, and code chunks they need.

Demo

asciicast

Requirements

- OpenAI API Key: Required for generating embeddings (local embedding support planned) - Git: Must be a git repository (respects .gitignore files) - Add .sourcerer/ to .gitignore: This directory stores the embedded vector database

Installation

Go

``shell go install github.com/st3v3nmw/sourcerer-mcp/cmd/sourcerer@latest `

Homebrew

`shell brew tap st3v3nmw/tap brew install st3v3nmw/tap/sourcerer `

Configuration

Claude Code

`shell claude mcp add sourcerer -e OPENAI_API_KEY=your-openai-api-key -e SOURCERER_WORKSPACE_ROOT=$(pwd) -- sourcerer `

mcp.json

`json { "mcpServers": { "sourcerer": { "command": "sourcerer", "env": { "OPENAI_API_KEY": "your-openai-api-key", "SOURCERER_WORKSPACE_ROOT": "/path/to/your/project" } } } } `

How it Works

Sourcerer builds a semantic search index of your codebase:

1. Code Parsing & Chunking

- Uses Tree-sitter to parse source files into ASTs - Extracts meaningful chunks (functions, classes, methods, types) with stable IDs - Each chunk includes source code, location info, and contextual summaries - Chunk IDs follow the pattern:
file.ext::TypeName::methodName

2. File System Integration

- Watches for file changes using
fsnotify - Respects .gitignore files via git check-ignore - Automatically re-indexes changed files - Stores metadata to track modification times

3. Vector Database

- Uses chromem-go for persistent vector storage in
.sourcerer/db/ - Generates embeddings via OpenAI's API for semantic similarity - Enables conceptual search rather than just text matching - Maintains chunks, their embeddings, and metadata

4. MCP Tools

-
semantic_search: Find code by concept/functionality - get_source_code: Retrieve specific chunks by ID - index_workspace: Manually trigger re-indexing - get_index_status`: Check indexing progress This approach allows AI agents to find relevant code without reading entire files, dramatically reducing token usage and cognitive load.

Supported Languages

Language support requires writing Tree-sitter queries to identify functions, classes, interfaces, and other code structures for each language. Supported: Go Planned: Python, TypeScript, JavaScript

Contributing

All contributions welcome!
No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.