Go Docs Mcp

by drolosoft

375 downloads Not rated yet
GitHub

About

Multi-format document MCP server — read, search, OCR, and extract from PDF, TXT, MD, DOCX, CSV, and images. Single Go binary, 12 tools, smart caching.

Explore

| Category | Tool | Description |
|----------|------|-------------|
| Discovery | list_documents | List all available documents with metadata (filename, format, page count, size) |
| Discovery | list_formats | List supported document formats and their dependencies |
| Reading | read_document | Read full text, a specific page, or page ranges from any supported document |
| Reading | read_url | Download a document from a URL and extract its text content |
| Reading | get_document_summary | Get the first 3 pages of text as a quick overview |
| Search | search_document | Case-insensitive full-text search with context and page hints |
| Analysis | get_document_metadata | Get full document metadata (title, author, dates, version, etc.) |
| Analysis | get_document_outline | Extract document outline / table of contents |
| Analysis | extract_tables | Extract tables from documents as structured data |
| Analysis | extract_images | Extract images from a document as base64-encoded data (max 10 per call) |
| OCR | ocr_document | Force OCR on a PDF — for scanned/image-based documents or garbled text |
| OCR | read_image | Extract text from an image file (PNG, JPG, TIFF) via OCR |

- Fast — mtime-based in-memory caching avoids redundant extraction
- Multi-format — PDF, TXT, MD, CSV, DOCX, and images from a single server
- OCR — automatic fallback to tesseract for image-based/scanned documents
- Secure — directory-locked access with path traversal prevention
- Simple — single binary, stdio transport, zero configuration required
- Portable — works on macOS and Linux

---

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Go Docs Mcp
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

- Go 1.25+ (install)
- poppler (provides pdftotext, pdfinfo, pdfimages, pdftoppm) — required for PDF support
- tesseract _(optional, for OCR support — scanned PDFs and images)_
- pandoc _(optional, for DOCX support)_

``bash

The server reads documents from a configured directory. Set DOCS_MCP_DIR to change it:

| Variable | Default | Description |
|----------|---------|-------------|
|
DOCS_MCP_DIR | ~/.docs-mcp/documents/ | Directory containing document files to serve |
|
PDF_MCP_DIR | _(backward compat alias)_ | Legacy alias — works the same as DOCS_MCP_DIR` |

Place your documents in the directory and the server will find them automatically. All supported formats (PDF, TXT, MD, CSV, DOCX, images) are detected.

---

Category

Tool

| Category | Tool | Description |
|----------|------|-------------|
| Discovery | list_documents | List all available documents with metadata (filename, format, page count, size) |
| Discovery | list_formats | List supported document formats and their dependencies |
| Reading | read_document | Read full text, a specific page, or page ranges from any supported document |
| Reading | read_url | Download a document from a URL and extract its text content |
| Reading | get_document_summary | Get the first 3 pages of text as a quick overview |
| Search | search_document | Case-insensitive full-text search with context and page hints |
| Analysis | get_document_metadata | Get full document metadata (title, author, dates, version, etc.) |
| Analysis | get_document_outline | Extract document outline / table of contents |
| Analysis | extract_tables | Extract tables from documents as structured data |
| Analysis | extract_images | Extract images from a document as base64-encoded data (max 10 per call) |
| OCR | ocr_document | Force OCR on a PDF — for scanned/image-based documents or garbled text |
| OCR | read_image | Extract text from an image file (PNG, JPG, TIFF) via OCR |

- Fast — mtime-based in-memory caching avoids redundant extraction
- Multi-format — PDF, TXT, MD, CSV, DOCX, and images from a single server
- OCR — automatic fallback to tesseract for image-based/scanned documents
- Secure — directory-locked access with path traversal prevention
- Simple — single binary, stdio transport, zero configuration required
- Portable — works on macOS and Linux

---

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "go docs mcp": {
            "docs": {
                "command": "go-docs-mcp",
                "env": {
                    "DOCS_MCP_DIR": "/path/to/your/documents"
                }
            }
        }
    }
}

McpServers

{
    "docs": {
        "command": "go-docs-mcp",
        "env": {
            "DOCS_MCP_DIR": "/path/to/your/documents"
        }
    }
}

🏆 Why go-docs-mcp?

| | Capability | go-docs-mcp | Node/TS | Python | Rust |
|--|-----------|:---:|:---:|:---:|:---:|
| ⚡ | Single binary, no runtime | ✅ | ❌ needs Node | ❌ needs Python | ✅ |
| 📦 | go install one-liner | ✅ | ❌ npm+deps | ❌ pip+venv | ❌ cargo |
| 📄 | Multi-format (6 types) | ✅ | ❌ one format | ❌ one format | ❌ one format |
| 🔍 | Full-text search | ✅ | ⚠️ partial | ✅ | ✅ |
| 👁️ | OCR (scanned PDFs) | ✅ | ❌ | ❌ | ⚠️ partial |
| 🖼️ | Image extraction | ✅ | ⚠️ partial | ✅ | ❌ |
| 📊 | Table extraction | ✅ | ⚠️ partial | ✅ | ✅ |
| 📑 | Document outline | ✅ | ❌ | ❌ | ⚠️ partial |
| 🌐 | Fetch from URL | ✅ | ⚠️ partial | ❌ | ❌ |
| 🔒 | Dir-locked, read-only | ✅ | ⚠️ varies | ⚠️ varies | ✅ |
| 💾 | Smart caching | ✅ | ❌ | ❌ | ❌ |
| 🏠 | Fully offline | ✅ | ✅ | ✅ | ✅ |

Every other document MCP server handles one format — a PDF server for PDFs, a DOCX server for DOCX. You'd need three tools to read three formats. go-docs-mcp reads them all from a single binary.

---

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.