Semcode

by GoodbyePlanet

421 downloads Not rated yet
GitHub

About

An MCP (Model Context Protocol) server providing hybrid semantic search over code across a set of GitHub repositories that you list in config.yaml. It parses symbols with Tree-sitter and indexes both code and git commit history, so AI clients can query them by natural language or

Explore

- Hybrid retrieval combining dense embeddings and BM25
- Incremental indexing—only changed files are re-embedded
- Supports 19 programming languages with framework-aware parsing
- Optional git commit history indexing with full diffs
- MCP tools for search, symbol lookup, and reindexing
- HTTP API for triggering index from CI/CD pipelines

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Semcode
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Prerequisites: Python 3.12+, Docker, GitHub token


uv sync

cp config.example.yaml config.yaml

Configure which repositories to index in config.yaml:

services:
  - name: my-service
    github_repo: owner/repo
    github_ref: main              # optional, defaults to "main" — branch, tag, or commit SHA
    root: src/main/java           # optional — limit indexing to this subdirectory (useful for monorepos)
    exclude: # optional — skip matching paths
      - "/vendor/"
      - "/node_modules/"

The indexer automatically discovers and indexes all files with recognised extensions. Use root to scope a service to a
subdirectory within a shared repo, and exclude to skip paths you don't want indexed (tests, build artifacts, generated
code, etc.).

| Variable | Default | Description |
|-----------------------------|-------------------------|-------------------------------------------------------------------------------|
| GITHUB_TOKEN | (required) | GitHub token with repo read access |
| QDRANT_URL | http://localhost:6333 | Qdrant connection URL |
| QDRANT_COLLECTION | code_symbols | Collection name for code symbol vectors |
| QDRANT_COMMITS_COLLECTION | git_commits | Collection name for commit message vectors |
| EMBEDDINGS_PROVIDER | jina | One of jina, jina-api, voyage, openai, ollama — see Embedding providers below |
| GIT_HISTORY_MAX_COMMITS | 500 | Max commits indexed per service |
| MCP_TRANSPORT | streamable-http | One of streamable-http, sse, stdio |
| MCP_HOST / MCP_PORT | 127.0.0.1 / 8090 | Server bind address |
| CONFIG_PATH | ./config.yaml | Path to the services config file |

search_code

Hybrid (dense + BM25) search by query, with optional filters for language, service, symbol type

find_symbol

Look up a symbol by name — exact match, or case-insensitive substring when `exact=false`

find_usages

Find code that references a given symbol name (semantic search, then excludes the definition itself)

get_code_context

Fetch the full source of a file — or a specific symbol within it — directly from GitHub

reindex

Trigger code indexing of one or all services (incremental by default; `force` to re-embed)

index_history

Index git commit history; automatically fetches diffs for commits missing them

search_commits

Search git commit history with natural language

get_commit

Get full details for a specific commit including changed files and diffs

list_indexed_services

List indexed services with chunk and file counts, languages, and last-indexed time

index_stats

Show Qdrant collection statistics and configured services

| Tool | Description |
|-------------------------|------------------------------------------------------------------------------------------------------|
| search_code | Hybrid (dense + BM25) search by query, with optional filters for language, service, symbol type |
| find_symbol | Look up a symbol by name — exact match, or case-insensitive substring when exact=false |
| find_usages | Find code that references a given symbol name (semantic search, then excludes the definition itself) |
| get_code_context | Fetch the full source of a file — or a specific symbol within it — directly from GitHub |
| reindex | Trigger code indexing of one or all services (incremental by default; force to re-embed) |
| index_history | Index git commit history; automatically fetches diffs for commits missing them |
| search_commits | Search git commit history with natural language |
| get_commit | Get full details for a specific commit including changed files and diffs |
| list_indexed_services | List indexed services with chunk and file counts, languages, and last-indexed time |
| index_stats | Show Qdrant collection statistics and configured services |

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "semcode": {
            "semcode": {
                "transport": "http",
                "url": "http://localhost:8090/mcp"
            }
        }
    }
}

McpServers

{
    "semcode": {
        "transport": "http",
        "url": "http://localhost:8090/mcp"
    }
}

An MCP (Model Context Protocol) server providing hybrid semantic search over code across a set of
GitHub repositories that you list in config.yaml. It parses symbols
with Tree-sitter and indexes both code and git commit history, so AI clients can query them by
natural language or by symbol name.

Hybrid retrieval combines dense embeddings with BM25, so both natural-language queries
("where do we publish order events?") and symbol-name lookups (PlaceOrderRequest) work well.

Submitted on mcpservers.org

How it works

1. Fetches source files from configured GitHub repositories
2. Parses code symbols (functions, classes, methods, components) using Tree-sitter
3. Generates two embeddings per symbol — a dense semantic vector (pluggable provider: Jina Code V2 by default, or
Voyage / OpenAI / Ollama) and a BM25 sparse vector keyed on code-identifier tokens (camelCase / snake_case split into
subwords)
4. Stores both in Qdrant and retrieves them with hybrid search — Reciprocal Rank Fusion (RRF) over the dense and
sparse results — so natural-language queries and symbol-name lookups both work well
5. Optionally indexes commit history into a separate Qdrant collection (dense-only)
6. Exposes search and indexing tools through the MCP protocol (and a small HTTP API)

Indexing is incremental — files are skipped when their Git blob SHA matches the last indexed version.
Files that no longer exist (or parse to zero symbols) are cleaned up automatically. Pass force: true
to re-embed everything.

Supported languages

Language is detected automatically from file extension or filename — no configuration needed.

Go, Java, Python, TypeScript / JavaScript (React), Rust, C#, C, C++, Ruby, PHP, Kotlin, Scala, Swift, Dart, Bash, SQL,
Lua, R, Dockerfile, Docker Compose, Markdown, JSON, HTML, CSS, XML.

Most parsers are framework-aware where it matters — Spring stereotypes and HTTP routes for Java/Kotlin, FastAPI/Pydantic
for Python, ASP.NET for C#, Rails for Ruby, Laravel/Symfony for PHP, React/SwiftUI/Flutter widgets, etc. See
server/parser/ for the per-language extraction details.

Setup

Prerequisites: Python 3.12+, Docker, GitHub token

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.