VisionAgent MCP Server

by landing-ai

292 downloads Not rated yet

About

VisionAgent MCP Server is a lightweight, side-car MCP server that runs locally on STDIN/STDOUT, translating tool calls from MCP‑compatible clients (Claude Desktop, Cursor, Cline) into authenticated HTTPS requests to Landing AI’s VisionAgent REST APIs. It enables natural‑language…

Explore

- Translates MCP tool calls into authenticated VisionAgent API requests.
- Supports agentic document analysis, text‑to‑object detection, text‑to‑instance segmentation, activity recognition, and depth estimation.
- Auto‑generates tool definitions from a live OpenAPI spec via npm run generate-tools.
- Renders masks, bounding boxes, and depth maps to files or inline previews.
- Validates arguments with Zod schemas derived from the OpenAPI spec.
- Outputs JSON and media results streamed back to the MCP client.

| Software | Minimum Version |
| ------------------------ | ---------------------------------------- |
| Node.js | 20 (LTS) |
| VisionAgent account | Any paid or free tier (needs API key) |
| MCP client | Claude Desktop / Cursor / Cline / etc. |

npm install -g vision-tools-mcp

{
"mcpServers": {
"VisionAgent": {
"command": "npx",
"args": ["vision-tools-mcp"],
"env": {
"VISION_AGENT_API_KEY": "<YOUR_API_KEY>",
"OUTPUT_DIRECTORY": "/path/to/output/directory",
"IMAGE_DISPLAY_ENABLED": "true" # or false, see below
}
}
}
}


4. Open your MCP-aware client.
5. Download street.png (from the assets folder in this directory, or you can choose any test image).
6. Paste the prompt below (or any prompt):


Detect all traffic lights in /path/to/mcp/vision-agent-mcp/assets/street.png

If your client supports inline resources, you’ll see bounding-box overlays; otherwise, the PNG is saved to your output directory, and the chat shows its path.

| ENV var | Required | Default | Purpose |
| ----------------------- | -------- | ---------- | ------------------------------------------------------ |
| VISION_AGENT_API_KEY | Yes | — | Landing AI auth token. |
| OUTPUT_DIRECTORY | No | — | Where rendered images / masks / depth maps are stored. |
| IMAGE_DISPLAY_ENABLED | No | true | false ➜ skip rendering |

1. Clone the repository:

bash
git clone https://github.com/landing-ai/vision-agent-mcp.git

2. Navigate into the project directory:

bash
cd vision-agent-mcp

3. Install dependencies:

bash
npm install

4. Build the project:

bash
npm run build

- VISION_AGENT_API_KEY - Required API key for VisionAgent authentication
- OUTPUT_DIRECTORY - Optional directory for saving processed outputs (supports relative and absolute paths)
- IMAGE_DISPLAY_ENABLED - Set to "true" to enable image visualization features

After building, configure your MCP client with the following settings:

json
{
"mcpServers": {
"VisionAgent": {
"command": "node",
"args": [
"/path/to/build/index.js"
],
"env": {
"VISION_AGENT_API_KEY": "<YOUR_API_KEY>",
"OUTPUT_DIRECTORY": "../../output",
"IMAGE_DISPLAY_ENABLED": "true"
}
}
}
}
``

> Note: Replace /path/to/build/index.js with the actual path to your built index.js` file, and set your environment variables as needed. For MCP clients without image display capabilities, like Cursor, set IMAGE_DISPLAY_ENABLED to False. For MCP clients with image display capabilities, like Claude Desktop, set IMAGE_DISPLAY_ENABLED to true to visualize tool outputs. Generally, MCP clients that support resources (see this list: https://modelcontextprotocol.io/clients) will support image display.

npm
build

> Beta – v0.1
> This project is early access and subject to breaking changes until v1.0.

VisionAgent MCP Server v0.1 - Overview

Modern LLM “agents” call external tools through the Model Context Protocol (MCP). VisionAgent MCP is a lightweight, side-car MCP server that runs locally on STDIN/STDOUT, translating each tool call from an MCP-compatible client (Claude Desktop, Cursor, Cline, etc.) into an authenticated HTTPS request to Landing AI’s VisionAgent REST APIs. The response JSON, plus any images or masks, is streamed back to the model so that you can issue natural-language computer-vision and document-analysis commands from your editor without writing custom REST code or loading an extra SDK.

📸 Demo

https://github.com/user-attachments/assets/2017fa01-0e7f-411c-a417-9f79562627b7

🧰 Supported Use Cases (v0.1)

| Capability | Description |
| ----------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| agentic-document-analysis | Parse PDFs / images to extract text, tables, charts, and diagrams taking into account layouts and other visual cues. Web Version here.|
| text-to-object-detection | Detect free-form prompts (“all traffic lights”) using OWLv2 / CountGD / Florence-2 / Agentic Object Detection (Web Version here); outputs bounding boxes. |
| text-to-instance-segmentation | Pixel-perfect masks via Florence-2 + Segment-Anything-v2 (SAM-2). |
| activity-recognition | Recognise multiple activities in video with start/end timestamps. |
| depth-pro | High-resolution monocular depth estimation for single images. |

> Run npm run generate-tools whenever VisionAgent releases new endpoints. The script fetches the latest OpenAPI spec and regenerates the local tool map automatically.

🗺 Table of Contents

1. Quick Start 2. Configuration 3. Example Prompts 4. Architecture & Flow 5. Developer Guide 6. Troubleshooting 7. Contributing 8. Security & Privacy

🚀 Quick Start

Get Your VisionAgent API Key

If you do not have a VisionAgent API key, create an account and obtain your API key.

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.