docs-mcp-server MCP Server
About
# docs-mcp-server MCP Server A MCP server for fetching and searching 3rd party package documentation. ## ✨ Key Features - 🌐 **Versatile Scraping:** Fetch documentation from diverse sources like websites, GitHub, npm, PyPI, or local files. - 🧠 **Intelligent Processing:** Automatically split content semantically and…
Details
- License
- MIT
Explore
- 🌐 Versatile Scraping: Fetch documentation from diverse sources like websites, GitHub, npm, PyPI, or local files.
- 🧠 Intelligent Processing: Automatically split content semantically and generate embeddings using your choice of models (OpenAI, Google Gemini, Azure OpenAI, AWS Bedrock, Ollama, and more).
- 💾 Optimized Storage: Leverage SQLite with sqlite-vec for efficient vector storage and FTS5 for robust full-text search.
- 🔍 Powerful Hybrid Search: Combine vector similarity and full-text search across different library versions for highly relevant results.
- ⚙️ Asynchronous Job Handling: Manage scraping and indexing tasks efficiently with a background job queue and MCP/CLI tools.
- 🐳 Simple Deployment: Get up and running quickly using Docker or npx.
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
docs-mcp-server MCP ServerCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
The following environment variables are supported to configure the embedding model behavior:
- DOCS_MCP_EMBEDDING_MODEL: Optional. Format: provider:model_name or just model_name (defaults to text-embedding-3-small). Supported providers and their required environment variables:
- openai (default): Uses OpenAI's embedding models
- OPENAI_API_KEY: Required. Your OpenAI API key
- OPENAI_ORG_ID: Optional. Your OpenAI Organization ID
- OPENAI_API_BASE: Optional. Custom base URL for OpenAI-compatible APIs (e.g., Ollama, Azure OpenAI)
- vertex: Uses Google Cloud Vertex AI embeddings
- GOOGLE_APPLICATION_CREDENTIALS: Required. Path to service account JSON key file
- gemini: Uses Google Generative AI (Gemini) embeddings
- GOOGLE_API_KEY: Required. Your Google API key
- aws: Uses AWS Bedrock embeddings
- AWS_ACCESS_KEY_ID: Required. AWS access key
- AWS_SECRET_ACCESS_KEY: Required. AWS secret key
- AWS_REGION or BEDROCK_AWS_REGION: Required. AWS region for Bedrock
- microsoft: Uses Azure OpenAI embeddings
- AZURE_OPENAI_API_KEY: Required. Azure OpenAI API key
- AZURE_OPENAI_API_INSTANCE_NAME: Required. Azure instance name
- AZURE_OPENAI_API_DEPLOYMENT_NAME: Required. Azure deployment name
- AZURE_OPENAI_API_VERSION: Required. Azure API version
There are two ways to run the docs-mcp-server:
This section covers running the server/CLI directly from the source code for development purposes. The primary usage method is now via the public Docker image as described in "Method 2".
This provides an isolated environment and exposes the server via HTTP endpoints.
1. Clone the repository:
git clone https://github.com/arabold/docs-mcp-server.git # Replace with actual URL if different
cd docs-mcp-server
2. Create
.env file:Copy the example and add your OpenAI key (see "Environment Setup" below).
cp .env.example .env
Note: This .env file setup is primarily needed when running the server from source or using the Docker method. When using the npx integration method, the OPENAI_API_KEY is set directly in the MCP configuration file.
1. Create a .env file based on .env.example:
bashcp .env.example .env
2. Update your OpenAI API key in .env:
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"docs-mcp-server mcp server": {
"MCP-DOC-Server-OpenRouter": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"\\"
]
}
}
}
}
McpServers
{
"MCP-DOC-Server-OpenRouter": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"\\"
]
}
}
A MCP server for fetching and searching 3rd party package documentation.
✨ Key Features
- 🌐 Versatile Scraping: Fetch documentation from diverse sources like websites, GitHub, npm, PyPI, or local files.
- 🧠 Intelligent Processing: Automatically split content semantically and generate embeddings using your choice of models (OpenAI, Google Gemini, Azure OpenAI, AWS Bedrock, Ollama, and more).
- 💾 Optimized Storage: Leverage SQLite with sqlite-vec for efficient vector storage and FTS5 for robust full-text search.
- 🔍 Powerful Hybrid Search: Combine vector similarity and full-text search across different library versions for highly relevant results.
- ⚙️ Asynchronous Job Handling: Manage scraping and indexing tasks efficiently with a background job queue and MCP/CLI tools.
- 🐳 Simple Deployment: Get up and running quickly using Docker or npx.
Overview
This project provides a Model Context Protocol (MCP) server designed to scrape, process, index, and search documentation for various software libraries and packages. It fetches content from specified URLs, splits it into meaningful chunks using semantic splitting techniques, generates vector embeddings using OpenAI, and stores the data in an SQLite database. The server utilizes sqlite-vec for efficient vector similarity search and FTS5 for full-text search capabilities, combining them for hybrid search results. It supports versioning, allowing documentation for different library versions (including unversioned content) to be stored and queried distinctly.
The server exposes MCP tools for:
- Starting a scraping job (scrape_docs): Returns a jobId immediately.
- Checking job status (get_job_status): Retrieves the current status and progress of a specific job.
- Listing active/completed jobs (list_jobs): Shows recent and ongoing jobs.
- Cancelling a job (cancel_job): Attempts to stop a running or queued job.
- Searching documentation (search_docs).
- Listing indexed libraries (list_libraries).
- Finding appropriate versions (find_version).
- Removing indexed documents (remove_docs).
- Fetching single URLs (fetch_url): Fetches a URL and returns its content as Markdown.
🆕 OpenRouter API 集成与多模型支持
Chat/Completions 功能
本服务已全面适配 OpenRouter API,支持主流大模型(GPT-4.1、Claude 3.7、Gemini 2.5、Grok、Qwen 等),并支持多模态输入(文本+图片)。
主要特性
- ✅ 支持 OpenRouter 官方所有主流模型,模型列表见src/utils/openrouter.ts 的 OPENROUTER_MODELS
- ✅ 支持多模态消息格式(如 text、image_url)
- ✅ 支持自定义 HTTP-Referer、X-Title 等 header,便于 openrouter.ai 统计和排名
- ✅ 支持 OpenRouter API 的所有扩展参数(如 stream、tools、temperature、max_tokens 等)
环境变量配置
-OPENAI_API_KEY:OpenRouter API Key(必填)
- OPENAI_API_BASE:OpenRouter API Base,推荐 https://openrouter.ai/api/v1
- MODEL_ID:默认模型(如 openai/gpt-4.1),可选
示例代码
import { openrouterChat } from './src/utils/openrouter';
const messages = [
{
role: 'user',
content: [
{ type: 'text', text: 'What is in this image?' },
{ type: 'image_url', image_url: { url: 'https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg' } }
]
}
];
const result = await openrouterChat({
model: 'openai/gpt-4.1',
messages,
referer: 'https://your-site.com', // 可选
xTitle: 'Your Site Name' // 可选
// 还可加 extraBody, headers 等参数
});
console.log(result);
支持的主流模型(部分示例)
- openai/gpt-4.1 - openai/gpt-4.1-mini - anthropic/claude-3.7-sonnet - google/gemini-2.5-pro-preview-03-25 - x-ai/grok-3-beta - qwen/qwen2.5-vl-32b-instruct:free - deepseek/deepseek-chat-v3-0324:free - thudm/glm-z1-32b:free - openrouter/auto - ...(详见源码 OPENROUTER_MODELS)更多 API 参数
如需支持流式输出、函数调用、system prompt、stop、temperature、max_tokens 等 OpenRouter API 参数,只需通过extraBody 字段传递即可,无需修改底层代码。
⚠️ Embedding 功能说明
> Embedding 功能已禁用!
>
> 本项目当前版本已彻底移除所有 embedding 相关实现和依赖,不再支持向量生成与检索。所有 embedding 相关 API 均会直接抛出异常提示。
>
> 仅保留全文检索与大模型 chat/completions 能力。
Configuration
The following environment variables are supported to configure the embedding model behavior:
Embedding Model Configuration
- DOCS_MCP_EMBEDDING_MODEL: Optional. Format: provider:model_name or just model_name (defaults to text-embedding-3-small). Supported providers and their required environment variables:
- openai (default): Uses OpenAI's embedding models
- OPENAI_API_KEY: Required. Your OpenAI API key
- OPENAI_ORG_ID: Optional. Your OpenAI Organization ID
- OPENAI_API_BASE: Optional. Custom base URL for OpenAI-compatible APIs (e.g., Ollama, Azure OpenAI)
- vertex: Uses Google Cloud Vertex AI embeddings
- GOOGLE_APPLICATION_CREDENTIALS: Required. Path to service account JSON key file
- gemini: Uses Google Generative AI (Gemini) embeddings
- GOOGLE_API_KEY: Required. Your Google API key
- aws: Uses AWS Bedrock embeddings
- AWS_ACCESS_KEY_ID: Required. AWS access key
- AWS_SECRET_ACCESS_KEY: Required. AWS secret key
- AWS_REGION or BEDROCK_AWS_REGION: Required. AWS region for Bedrock
- microsoft: Uses Azure OpenAI embeddings
- AZURE_OPENAI_API_KEY: Required. Azure OpenAI API key
- AZURE_OPENAI_API_INSTANCE_NAME: Required. Azure instance name
- AZURE_OPENAI_API_DEPLOYMENT_NAME: Required. Azure deployment name
- AZURE_OPENAI_API_VERSION: Required. Azure API version
Vector Dimensions
The database schema uses a fixed dimension of 1536 for embedding vectors. Only models that produce vectors with dimension ≤ 1536 are supported, except for certain providers (like Gemini) that support dimension reduction.
For OpenAI-compatible APIs (like Ollama), use the openai provider with OPENAI_API_BASE pointing to your endpoint.
These variables can be set regardless of how you run the server (Docker, npx, or from source).
Running the MCP Server
There are two ways to run the docs-mcp-server:
Option 1: Using Docker (Recommended)
This is the recommended approach for most users. It's easy, straightforward, and doesn't require Node.js to be installed.
1. Ensure Docker is installed and running.
2. Configure your MCP settings:
Claude/Cline/Roo Configuration Example:
Add the following configuration block to your MCP settings file (adjust path as needed):
{
"mcpServers": {
"docs-mcp-server": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"-e",
"OPENAI_API_KEY",
"-v",
"docs-mcp-data:/data",
"ghcr.io/arabold/docs-mcp-server:latest"
],
"env": {
"OPENAI_API_KEY": "sk-proj-..." // Required: Replace with your key
},
"disabled": false,
"autoApprove": []
}
}
}
Remember to replace "sk-proj-..." with your actual OpenAI API key and restart the application.
3. That's it! The server will now be available to your AI assistant.
Docker Container Settings:
- -i: Keep STDIN open, crucial for MCP communication over stdio.
- --rm: Automatically remove the container when it exits.
- -e OPENAI_API_KEY: Required. Set your OpenAI API key.
- -v docs-mcp-data:/data: Required for persistence. Mounts a Docker named volume docs-mcp-data to store the database. You can replace with a specific host path if preferred (e.g., -v /path/on/host:/data).
Any of the configuration environment variables (see Configuration above) can be passed to the container using the -e flag. For example:
```bash
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



