VisionAgent MCP Server
About
VisionAgent MCP Server is a lightweight, side-car MCP server that runs locally on STDIN/STDOUT, translating tool calls from MCP‑compatible clients (Claude Desktop, Cursor, Cline) into authenticated HTTPS requests to Landing AI’s VisionAgent REST APIs. It enables natural‑language…
Explore
- Translates MCP tool calls into authenticated VisionAgent API requests.
- Supports agentic document analysis, text‑to‑object detection, text‑to‑instance segmentation, activity recognition, and depth estimation.
- Auto‑generates tool definitions from a live OpenAPI spec via npm run generate-tools.
- Renders masks, bounding boxes, and depth maps to files or inline previews.
- Validates arguments with Zod schemas derived from the OpenAPI spec.
- Outputs JSON and media results streamed back to the MCP client.
| Software | Minimum Version |
| ------------------------ | ---------------------------------------- |
| Node.js | 20 (LTS) |
| VisionAgent account | Any paid or free tier (needs API key) |
| MCP client | Claude Desktop / Cursor / Cline / etc. |
npm install -g vision-tools-mcp
{
"mcpServers": {
"VisionAgent": {
"command": "npx",
"args": ["vision-tools-mcp"],
"env": {
"VISION_AGENT_API_KEY": "<YOUR_API_KEY>",
"OUTPUT_DIRECTORY": "/path/to/output/directory",
"IMAGE_DISPLAY_ENABLED": "true" # or false, see below
}
}
}
}
4. Open your MCP-aware client.
5. Download street.png (from the assets folder in this directory, or you can choose any test image).
6. Paste the prompt below (or any prompt):
Detect all traffic lights in /path/to/mcp/vision-agent-mcp/assets/street.png
If your client supports inline resources, you’ll see bounding-box overlays; otherwise, the PNG is saved to your output directory, and the chat shows its path.
| ENV var | Required | Default | Purpose |
| ----------------------- | -------- | ---------- | ------------------------------------------------------ |
| VISION_AGENT_API_KEY | Yes | — | Landing AI auth token. |
| OUTPUT_DIRECTORY | No | — | Where rendered images / masks / depth maps are stored. |
| IMAGE_DISPLAY_ENABLED | No | true | false ➜ skip rendering |
1. Clone the repository:
bashgit clone https://github.com/landing-ai/vision-agent-mcp.git
2. Navigate into the project directory:
bashcd vision-agent-mcp
3. Install dependencies:
bashnpm install
4. Build the project:
bashnpm run build
- VISION_AGENT_API_KEY - Required API key for VisionAgent authentication
- OUTPUT_DIRECTORY - Optional directory for saving processed outputs (supports relative and absolute paths)
- IMAGE_DISPLAY_ENABLED - Set to "true" to enable image visualization features
After building, configure your MCP client with the following settings:
json{
"mcpServers": {
"VisionAgent": {
"command": "node",
"args": [
"/path/to/build/index.js"
],
"env": {
"VISION_AGENT_API_KEY": "<YOUR_API_KEY>",
"OUTPUT_DIRECTORY": "../../output",
"IMAGE_DISPLAY_ENABLED": "true"
}
}
}
}
``
> Note: Replace
/path/to/build/index.js with the actual path to your built index.js` file, and set your environment variables as needed. For MCP clients without image display capabilities, like Cursor, set IMAGE_DISPLAY_ENABLED to False. For MCP clients with image display capabilities, like Claude Desktop, set IMAGE_DISPLAY_ENABLED to true to visualize tool outputs. Generally, MCP clients that support resources (see this list: https://modelcontextprotocol.io/clients) will support image display.> Beta – v0.1
> This project is early access and subject to breaking changes until v1.0.
VisionAgent MCP Server v0.1 - Overview
Modern LLM “agents” call external tools through the Model Context Protocol (MCP). VisionAgent MCP is a lightweight, side-car MCP server that runs locally on STDIN/STDOUT, translating each tool call from an MCP-compatible client (Claude Desktop, Cursor, Cline, etc.) into an authenticated HTTPS request to Landing AI’s VisionAgent REST APIs. The response JSON, plus any images or masks, is streamed back to the model so that you can issue natural-language computer-vision and document-analysis commands from your editor without writing custom REST code or loading an extra SDK.
📸 Demo
https://github.com/user-attachments/assets/2017fa01-0e7f-411c-a417-9f79562627b7
🧰 Supported Use Cases (v0.1)
| Capability | Description |
| ----------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| agentic-document-analysis | Parse PDFs / images to extract text, tables, charts, and diagrams taking into account layouts and other visual cues. Web Version here.|
| text-to-object-detection | Detect free-form prompts (“all traffic lights”) using OWLv2 / CountGD / Florence-2 / Agentic Object Detection (Web Version here); outputs bounding boxes. |
| text-to-instance-segmentation | Pixel-perfect masks via Florence-2 + Segment-Anything-v2 (SAM-2). |
| activity-recognition | Recognise multiple activities in video with start/end timestamps. |
| depth-pro | High-resolution monocular depth estimation for single images. |
> Run npm run generate-tools whenever VisionAgent releases new endpoints. The script fetches the latest OpenAPI spec and regenerates the local tool map automatically.
🗺 Table of Contents
1. Quick Start 2. Configuration 3. Example Prompts 4. Architecture & Flow 5. Developer Guide 6. Troubleshooting 7. Contributing 8. Security & Privacy🚀 Quick Start
Get Your VisionAgent API Key
If you do not have a VisionAgent API key, create an account and obtain your API key.```bash
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



