Docsearch Mcp

by PatrickKoss

278 downloads Not rated yet

About

A local‑first document search and indexing system that provides hybrid semantic + keyword search across local files (including PDFs) and Confluence pages through the Model Context Protocol (MCP). It is designed for AI assistants like Claude Code/Desktop to access documentation…

Explore

- 🔍 Hybrid Search: Combines full-text search (FTS) with vector similarity for optimal results
- 📁 Multi-Source: Index local files (code, docs, PDFs) and Confluence spaces
- 📄 PDF Support: Extract and search text from PDF documents with metadata preservation
- 🖼️ Image Search: AI-powered image description and search for diagrams, screenshots, and charts
- 🗄️ Database Flexibility: Support for SQLite (local-first) and PostgreSQL (scalable)
- 🤖 MCP Integration: Seamless integration with Claude Code and other MCP-compatible tools
- 💻 CLI Tool: Standalone command-line interface with multiple output formats
- ⚡ Real-time Updates: File watching with automatic re-indexing
- 🎯 Smart Chunking: Intelligent text chunking for code, documentation, and PDFs
- 📊 Multiple Output Formats: Text, JSON, and YAML output for search results
- 🔒 Secure: API keys and sensitive data stay on your machine

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Docsearch Mcp
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

npm install -g docsearch-mcp

Add to your MCP client configuration:

{
  "mcpServers": {
    "docsearch": {
      "command": "npx",
      "args": ["docsearch-mcp", "start"]
    }
  }
}

For development or customization:


make setup

pnpm install
cp .env.example .env

docsearch-mcp start


bash

When using the official Docker image, create a .env file:

echo "OPENAI_API_KEY=your-openai-key" > .env
echo "FILE_ROOTS=/app/documents" >> .env

Key considerations for Docker deployment:

- Document Volume: Mount your documents directory to /app/documents in the container
- Data Persistence: Mount your data directory to /app/data to persist the SQLite database
- User Permissions: Use --user $(id -u):$(id -g) to avoid permission issues with mounted volumes
- Environment Variables: Pass configuration via .env file or environment variables
- Official Image: Use ghcr.io/patrickkoss/docsearch-mcp:v0.0.2 for production deployments

docker run --rm -i \
-v /absolute/path/to/documents:/app/documents \
-v /absolute/path/to/data:/app/data \
--env-file .env \
--user $(id -u):$(id -g) \
ghcr.io/patrickkoss/docsearch-mcp:v0.0.5

Create a .env file from .env.example and configure:


docsearch ingest files
docsearch ingest confluence
docsearch ingest all --watch

docsearch ingest files --file-roots "./src,./docs" --openai-api-key sk-xxx

docsearch --config prod.env --db-path /data/prod.db ingest all

The CLI supports multiple configuration sources in order of precedence:

1. Command-line arguments (highest priority)
2. Custom config file (--config path/to/.env)
3. .env.local file
4. .env file (lowest priority)

All configuration options are passed via environment variables to the MCP server:

| Environment Variable | Default | Description |
| --------------------- | -------- | ----------------------------------------- |
| EMBEDDINGS_PROVIDER | openai | Embeddings provider (openai or tei) |
| OPENAI_API_KEY | - | OpenAI API key (required if using OpenAI) |

| Environment Variable | Default | Description |
| ---------------------------- | ------------------------------------------------------------- | ---------------------------------------- |
| OPENAI_BASE_URL | - | OpenAI base URL (for custom endpoints) |
| OPENAI_EMBED_MODEL | text-embedding-3-small | OpenAI embedding model |
| OPENAI_EMBED_DIM | 1536 | OpenAI embedding dimension |
| TEI_ENDPOINT | - | Text Embeddings Inference endpoint |
| FILE_ROOTS | . | File roots to index (comma-separated) |
| FILE_INCLUDE_GLOBS | /.{go,ts,tsx,js,py,rs,java,md,mdx,txt,yaml,yml,json,pdf} | File include patterns |
| FILE_EXCLUDE_GLOBS |
/{.git,node_modules,dist,build,target}/ | File exclude patterns |
| CONFLUENCE_BASE_URL | - | Confluence base URL |
| CONFLUENCE_EMAIL | - | Confluence email |
| CONFLUENCE_API_TOKEN | - | Confluence API token |
| CONFLUENCE_SPACES | - | Confluence spaces (comma-separated) |
| DB_TYPE | sqlite | Database type (sqlite or postgresql) |
| DB_PATH | ./data/index.db | SQLite database path |
| POSTGRES_CONNECTION_STRING | - | PostgreSQL connection string |

1. Install the package:

bash
npm install -g docsearch-mcp

2.
Create your configuration:

bash

{
  "mcpServers": {
    "docsearch-work": {
      "command": "npx",
      "args": ["docsearch-mcp", "start"],
      "env": {
        "OPENAI_API_KEY": "sk-your-key",
        "FILE_ROOTS": "/work/projects/frontend,/work/projects/backend",
        "FILE_INCLUDE_GLOBS": "/.{ts,tsx,js,py,md,yaml}",
        "DB_PATH": "/work/data/work-index.db",
        "CONFLUENCE_BASE_URL": "https://company.atlassian.net",
        "CONFLUENCE_SPACES": "DEV,API,DOCS"
      }
    },
    "docsearch-personal": {
      "command": "npx",
      "args": ["docsearch-mcp", "start"],
      "env": {
        "OPENAI_API_KEY": "sk-your-key",
        "FILE_ROOTS": "/home/user/projects,/home/user/documents",
        "DB_PATH": "/home/user/.docsearch/personal.db",
        "EMBEDDINGS_PROVIDER": "tei",
        "TEI_ENDPOINT": "http://localhost:8080/embeddings"
      }
    }
  }
}
{
  "mcpServers": {
    "docsearch": {
      "command": "npx",
      "args": ["docsearch-mcp", "start"],
      "env": {
        "OPENAI_API_KEY": "sk-your-key",
        "DB_TYPE": "postgresql",
        "POSTGRES_CONNECTION_STRING": "postgresql://user:pass@localhost:5432/docsearch",
        "FILE_ROOTS": ".,/other/projects"
      }
    }
  }
}

npm install -g docsearch-mcp

echo "OPENAI_API_KEY=sk-your-key-here" > .env
echo "FILE_ROOTS=.,../other-project" >> .env

echo "OPENAI_API_KEY=sk-your-key" > .env
echo "FILE_ROOTS=/app/documents" >> .env

Team Documentation Hub:

json
{
"mcpServers": {
"team-knowledge": {
"command": "npx",
"args": ["docsearch-mcp", "start"],
"env": {
"OPENAI_API_KEY": "sk-your-key",
"FILE_ROOTS": "/team/frontend,/team/backend,/team/mobile,/team/docs",
"CONFLUENCE_BASE_URL": "https://company.atlassian.net",
"CONFLUENCE_SPACES": "TEAM,API,ARCH,DEPLOY",
"CONFLUENCE_EMAIL": "[email protected]",
"CONFLUENCE_API_TOKEN": "your-token",
"DB_TYPE": "postgresql",
"POSTGRES_CONNECTION_STRING": "postgresql://user:pass@localhost/team_docs"
}
}
}
}

Personal Research Setup:

bash

docsearch-mcp ingest files \
--file-roots "/research/papers,/research/reports" \
--file-include-globs "*/.{pdf,docx,md,txt}" \
--openai-embed-model "text-embedding-3-large" \
--openai-embed-dim 3072

docsearch-mcp search "kubernetes deployment" \
--source confluence \
--repo DEVOPS \
--mode keyword

The MCP server provides these tools for Claude Code:

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "docsearch mcp": {
            "docsearch": {
                "command": "npx",
                "args": [
                    "docsearch-mcp",
                    "start"
                ],
                "env": {
                    "OPENAI_API_KEY": "your-openai-key",
                    "EMBEDDINGS_PROVIDER": "openai",
                    "FILE_ROOTS": ".,../other-project",
                    "DB_PATH": "/path/to/your/index.db"
                }
            }
        }
    }
}

McpServers

{
    "docsearch": {
        "command": "npx",
        "args": [
            "docsearch-mcp",
            "start"
        ],
        "env": {
            "OPENAI_API_KEY": "your-openai-key",
            "EMBEDDINGS_PROVIDER": "openai",
            "FILE_ROOTS": ".,../other-project",
            "DB_PATH": "/path/to/your/index.db"
        }
    }
}

TypeScript
Node.js
MCP

A local-first document search and indexing system that provides hybrid semantic + keyword search across local files (including PDFs) and Confluence pages through the Model Context Protocol (MCP). Perfect for AI assistants like Claude Code/Desktop to access your documentation, codebase, and research materials.

✨ Features

- 🔍 Hybrid Search: Combines full-text search (FTS) with vector similarity for optimal results
- 📁 Multi-Source: Index local files (code, docs, PDFs) and Confluence spaces
- 📄 PDF Support: Extract and search text from PDF documents with metadata preservation
- 🖼️ Image Search: AI-powered image description and search for diagrams, screenshots, and charts
- 🗄️ Database Flexibility: Support for SQLite (local-first) and PostgreSQL (scalable)
- 🤖 MCP Integration: Seamless integration with Claude Code and other MCP-compatible tools
- 💻 CLI Tool: Standalone command-line interface with multiple output formats
- ⚡ Real-time Updates: File watching with automatic re-indexing
- 🎯 Smart Chunking: Intelligent text chunking for code, documentation, and PDFs
- 📊 Multiple Output Formats: Text, JSON, and YAML output for search results
- 🔒 Secure: API keys and sensitive data stay on your machine

🚀 Installation & Usage

npm Package

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.