Clawd Cursor
About
The local MCP server that gives any AI agent safe desktop control. Drives Windows, macOS, and Linux GUIs through accessibility, OCR, or vision fallback. 6 compact compound tools covering 97 primitives. Model-agnostic. Local-only by default; bearer-token auth on HTTP; every tool c
Explore
- Two MCP transports: stdio for editors, HTTP for daemons.
- Six compound tools (Anthropic computer_20250124-style) plus granular tools.
- Cheapest-tier-first pipeline: accessibility → OCR → screenshot → vision.
- Ground-truth verification with six independent signals and weighted voting.
- Reflector loop that emits structured failure causes and suggested strategies.
- Single safety chokepoint (safety.evaluate()) on every tool call.
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
Clawd CursorCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
Sixty seconds from zero to a tool-calling agent on your desktop.
Pick your mode first:
| Your situation | Use | Why |
|---|---|---|
| AI lives in your editor (Claude Code, Cursor, Windsurf, Zed) | clawdcursor mcp | stdio MCP server. Editor spawns it on demand. No daemon, no port. |
| You're building an agent that runs unattended | clawdcursor agent | HTTP MCP daemon on 127.0.0.1:3847. Has its own LLM brain optionally configured via doctor. |
| Your agent has its own brain — you just want the tools as an HTTP endpoint | clawdcursor agent --no-llm | Same daemon, no built-in pipeline, no scheduler startup, no credential validation. Pure tool surface. |
Windows (PowerShell):
irm https://clawdcursor.com/install.ps1 | iex
macOS / Linux:
curl -fsSL https://clawdcursor.com/install.sh | bash
Then verify and configure:
clawdcursor --version # smoke-test the install
clawdcursor consent --accept # one-time desktop-control consent (required)
clawdcursor status # cross-check permissions + AI config
clawdcursor doctor # (optional) configure an LLM provider end-to-end
clawdcursor agent # OR clawdcursor mcp — see the table above
The installer clones into ~/clawdcursor, runs npm install, builds, and npm links a global shim. Runtime state lives at ~/.clawdcursor/ (auth token, pidfiles, logs). It does not edit any agent host config — that step is below.
Wire it into Claude Code, Cursor, Windsurf, or Zed:
// ~/.claude/settings.json (or your editor's MCP config)
{
"mcpServers": {
"clawdcursor": {
"command": "clawdcursor",
"args": ["mcp", "--compact"]
}
}
}
That's it. Ask your agent to "open Outlook and reply to the latest email from Sarah" and watch it run.
> macOS: run clawdcursor grant to walk through Accessibility + Screen Recording permissions.
> Linux: install tesseract-ocr, python3-gi, gir1.2-atspi-2.0, and (Wayland only) ydotool or wtype.
---
clawdcursor consent Manage desktop-control consent (--accept / --revoke / --status)
clawdcursor grant Grant macOS permissions (interactive, macOS only)
clawdcursor doctor Verify permissions, configure AI provider + models
clawdcursor status Readiness check (consent, permissions, AI config)
computer
`screenshot`, `click`, `double_click`, `right_click`, `triple_click`, `hover`, `scroll`, `scroll_horizontal`, `drag`, `drag_path`, `type`, `key`, `wait`
accessibility
`read_tree`, `find`, `get_element`, `focused`, `invoke`, `focus`, `set_value`, `get_value`, `expand`, `collapse`, `toggle`, `select`, `state`, `list_children`, `wait_for`
window
`list`, `active`, `focus`, `maximize`, `minimize`, `restore`, `close`, `resize`, `list_displays`, `screen_size`, `open_app`, `open_file`, `open_url`, `switch_tab`, `navigate`
system
`clipboard_read`, `clipboard_write`, `system_time`, `ocr`, `undo`, `shortcuts_list`, `shortcuts_run`, `delegate`, `detect_webview`, `relaunch_with_cdp`, `app_guide`, `detect_app`, `classify_task`, `system_prompt`
browser
`connect`, `page_context`, `read_text`, `click`, `type`, `select_option`, `evaluate`, `wait_for`, `list_tabs`, `switch_tab`, `scroll`
task
`{instruction: string}` — hand off the whole task to the pipeline. No `action` enum.
curl -s -X POST http://127.0.0.1:3847/mcp \
-H "Authorization: Bearer $(cat ~/.clawdcursor/token)" \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
---
Two catalogs, side by side. Agents pick the shape that fits.
Anthropic computer_20250124-style: one tool per capability, an action enum for the verb. The compact catalog is roughly an order of magnitude smaller than the granular surface, which keeps small models (Haiku, Kimi, Ollama) focused on the action choice instead of drowning in primitives. Default for every agent that doesn't explicitly need one schema per primitive.
| Tool | Most-used actions |
|---|---|
| computer | screenshot, click, double_click, right_click, triple_click, hover, scroll, scroll_horizontal, drag, drag_path, type, key, wait |
| accessibility | read_tree, find, get_element, focused, invoke, focus, set_value, get_value, expand, collapse, toggle, select, state, list_children, wait_for |
| window | list, active, focus, maximize, minimize, restore, close, resize, list_displays, screen_size, open_app, open_file, open_url, switch_tab, navigate |
| system | clipboard_read, clipboard_write, system_time, ocr, undo, shortcuts_list, shortcuts_run, delegate, detect_webview, relaunch_with_cdp, app_guide, detect_app, classify_task, system_prompt |
| browser | connect, page_context, read_text, click, type, select_option, evaluate, wait_for, list_tabs, switch_tab, scroll |
| task | {instruction: string} — hand off the whole task to the pipeline. No action enum. |
One schema per verb. Use this when your runtime requires every primitive as a top-level tool. The full catalog is visible through MCP tools/list on either transport.
A typical turn:
js// Compact — recommended
computer({ action: "key", combo: "mod+s" }) // resolves to Cmd+S / Ctrl+S
accessibility({ action: "invoke", name: "Send" })
window({ action: "open_app", name: "Outlook" })
system({ action: "ocr" }) // OS-level OCR, no LLM vision
task({ instruction: "open Notepad and type hello" }) // full pipeline
```
---
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"clawd cursor": {
"clawdcursor": {
"command": "clawdcursor",
"args": [
"mcp",
"--compact"
]
}
}
}
}
McpServers
{
"clawdcursor": {
"command": "clawdcursor",
"args": [
"mcp",
"--compact"
]
}
}
- Works where APIs don't exist. Native apps. Legacy enterprise tools. Web portals behind SSO that block headless browsers. Anything inside Citrix or RDP. If pixels reach the screen, your agent can drive it.
- Model-agnostic. Claude, GPT, Gemini, Llama, Kimi, anything local via Ollama — any tool-calling LLM. Text and vision can be different models from different vendors.
- App-agnostic. No per-app plugins, no per-service auth. The same six compound tools drive Outlook, Figma, your bank, and that 2003-era ERP.
- Cheapest-tier-first pipeline (for natural-language tasks). When you hand a whole instruction to the autonomous agent via task({...}) / submit_task, the planner starts at accessibility (free), escalates to OCR (cheap), then screenshot (medium), then vision (expensive). The Reflector feeds verifier signals back so the loop doesn't keep paying for vision when text would work. Direct tool calls (when your own agent loop drives the compound or granular tools) skip the pipeline and go straight to the requested tool — you bring the planning, clawdcursor brings the primitives.
- Local-only by default. Server binds to 127.0.0.1. Screenshots stay in RAM unless you point a cloud model at them. No telemetry.
- One protocol, two transports. MCP over stdio for editor hosts; MCP over HTTP for daemons. Same tool catalog, same JSON-RPC envelope.
When NOT to use it
Clawd Cursor is GUI control. It's slower than an API, less reliable than a script, and burns more tokens than a direct file edit. If a better path exists, take it:
| Better option | When |
|---|---|
| Native API (Gmail API, GitHub API, Stripe API, …) | The service has one. Use it. |
| CLI (git, gh, aws, npm, curl, sqlite3) | The work fits a shell tool. Use it. |
| Direct file edit | The data lives in a file you can write. Edit it. |
| Browser automation already wired up (Playwright, Puppeteer) for this exact site | Faster, more deterministic. Use it. |
Reach for Clawd Cursor when none of those apply — the legacy ERP with no REST, the Electron app whose dialog can't be scripted, the Excel macro behind a Citrix session. The pipeline pays the cheap costs first; if structured paths work it never escalates to vision. But the design point is the last mile, not the first call.
(SKILL.md enforces this as a hard 4-gate rule for AI agents calling the tool surface. Humans get the softer table above.)
---
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



