AnyCrawl
About
[AnyCrawl](https://anycrawl.dev) MCP Server, Powerful web scraping and crawling for Cursor, Claude, and other LLM clients via the Model Context Protocol (MCP).
Explore
- Web Scraping: Extract content from single URLs with multiple output formats
- Website Crawling: Crawl entire websites with configurable depth and limits
- Search Engine Integration: Search the web and optionally scrape results
- Multiple Engines: Auto mode (intelligent selection), Playwright, Cheerio, and Puppeteer
- Flexible Output: Markdown, HTML, text, screenshots, and structured JSON
- Async Operations: Non-blocking crawl jobs with status monitoring
- Error Handling: Robust error handling and logging
- Multiple Modes: STDIO (default), MCP(HTTP), SSE; cloud-ready with Nginx proxy
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
AnyCrawlCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
AnyCrawl MCP Server supports two deployment modes: Cloud Service (recommended) and Self-Hosted.
ANYCRAWL_API_KEY=YOUR-API-KEY npx -y anycrawl-mcp
npm install -g anycrawl-mcp-server
ANYCRAWL_API_KEY=YOUR-API-KEY anycrawl-mcp
For self-hosted deployments, configure the base URL to point to your own AnyCrawl API instance:
export ANYCRAWL_API_KEY="your-api-key-here"
export ANYCRAWL_BASE_URL="https://your-api-server.com" # Your self-hosted API URL
For local development with custom host/port:
export ANYCRAWL_HOST="127.0.0.1" # Default: mcp.anycrawl.dev (cloud)
export ANYCRAWL_PORT="3000" # Default: 3000
Configuring Cursor. Note: Requires Cursor v0.45.6+.
For Cursor v0.48.6 and newer, add this to your MCP Servers settings:
json{
"mcpServers": {
"anycrawl-mcp": {
"command": "npx",
"args": ["-y", "anycrawl-mcp"],
"env": {
"ANYCRAWL_API_KEY": "YOUR-API-KEY"
}
}
}
}
For Cursor v0.45.6:
1. Open Cursor Settings → Features → MCP Servers → "+ Add New MCP Server"
2. Name: "anycrawl-mcp" (or your preferred name)
3. Type: "command"
4. Command:
bashenv ANYCRAWL_API_KEY=YOUR-API-KEY npx -y anycrawl-mcp
On Windows, if you encounter issues:
bashcmd /c "set ANYCRAWL_API_KEY=YOUR-API-KEY && npx -y anycrawl-mcp"
For manual installation, add this JSON to your User Settings (JSON) in VS Code (Command Palette → Preferences: Open User Settings (JSON)):
json{
"mcp": {
"inputs": [
{
"type": "promptString",
"id": "apiKey",
"description": "AnyCrawl API Key",
"password": true
}
],
"servers": {
"anycrawl": {
"command": "npx",
"args": ["-y", "anycrawl-mcp"],
"env": {
"ANYCRAWL_API_KEY": "${input:apiKey}"
}
}
}
}
}
Optionally, place the following in .vscode/mcp.json in your workspace to share config:
json{
"inputs": [
{
"type": "promptString",
"id": "apiKey",
"description": "AnyCrawl API Key",
"password": true
}
],
"servers": {
"anycrawl": {
"command": "npx",
"args": ["-y", "anycrawl-mcp"],
"env": {
"ANYCRAWL_API_KEY": "${input:apiKey}"
}
}
}
}
Add this to ~/.codeium/windsurf/model_config.json:
json{
"mcpServers": {
"mcp-server-anycrawl": {
"command": "npx",
"args": ["-y", "anycrawl-mcp"],
"env": {
"ANYCRAWL_API_KEY": "YOUR_API_KEY"
}
}
}
}
The SSE (Server-Sent Events) mode provides a web-based interface for MCP communication, ideal for web applications, testing, and integration with web-based LLM clients.
Optional server settings for local/self-hosted deployments:
bashexport ANYCRAWL_PORT=3000 # Default: 3000
export ANYCRAWL_HOST=127.0.0.1 # Set to override cloud default (mcp.anycrawl.dev)
For other MCP/SSE clients that support SSE transport, use this configuration:
json{
"mcpServers": {
"anycrawl": {
"type": "sse",
"url": "https://mcp.anycrawl.dev/{API_KEY}/sse",
"name": "AnyCrawl MCP Server",
"description": "Web scraping and crawling tools"
}
}
}
or
json{
"mcpServers": {
"AnyCrawl": {
"type": "streamable_http",
"url": "https://mcp.anycrawl.dev/{API_KEY}/mcp"
}
}
}
Environment Setup:
bash
Configure Cursor to connect to your HTTP MCP server.
Local HTTP Streamable Server:
{
"mcpServers": {
"anycrawl-http-local": {
"type": "streamable_http",
"url": "http://127.0.0.1:3000/mcp"
}
}
}
Cloud HTTP Streamable Server:
{
"mcpServers": {
"anycrawl-http-cloud": {
"type": "streamable_http",
"url": "https://mcp.anycrawl.dev/{API_KEY}/mcp"
}
}
}
Note: For HTTP modes, set ANYCRAWL_API_KEY (and optional host/port) in the server process environment or in the URL. Cursor does not need your API key when using streamable_http.
anycrawl_scrape
Scrape a single URL and extract content in selected formats. Best for: One known page (articles, docs, product pages). Not recommended for: Multi-page coverage (use anycrawl_crawl) or open-ended discovery (use anycrawl_search). RECOMMENDED: Use 'playwright' engine for best results with dynamic content and modern websites. Usage (parameters): - url: HTTP/HTTPS URL to scrape (string, required) - engine: 'playwright' | 'cheerio' | 'puppeteer' (required, default: 'playwright') - proxy: Proxy URL (string, optional) - formats: Output formats ['markdown'|'html'|'text'|'screenshot'|'screenshot@fullPage'|'rawHtml'|'json'] (optional) - timeout: Request timeout in ms (number, optional) - retry: Enable auto-retry on failure (boolean, optional) - wait_for: Wait in ms for dynamic pages (number, optional) - include_tags: HTML tags to include (string[], optional) - exclude_tags: HTML tags to exclude (string[], optional) - json_options: { schema?, user_prompt?, schema_name?, schema_description? } (optional) - extract_source: 'html' | 'markdown' (optional) Returns: { url, status, jobId?, title?, html?, markdown?, metadata?, timestamp? } Examples: - Recommended: { "url": "https://example.com", "engine": "playwright" } - With JSON extraction: { "url": "https://news.ycombinator.com", "engine": "playwright", "formats": ["markdown"], "json_options": { "user_prompt": "Extract titles", "schema_name": "Articles" } } - With JSON schema extraction: { "url": "https://example.com/article", "engine": "playwright", "json_options": { "schema_name": "Article", "schema_description": "Extract article metadata and content", "schema": { "type": "object", "properties": { "title": { "type": "string" }, "author": { "type": "string" }, "date": { "type": "string" }, "content": { "type": "string" } }, "required": ["title", "content"] } } }
anycrawl_crawl
Crawl an entire website with configurable depth and limits. Best for: Multi-page coverage, site mapping, content discovery. Not recommended for: Single pages (use anycrawl_scrape) or open-ended discovery (use anycrawl_search). RECOMMENDED: Use 'playwright' engine for best results with dynamic content and modern websites. Usage (parameters): - url: Starting URL to crawl (string, required) - engine: 'playwright' | 'cheerio' | 'puppeteer' (required, default: 'playwright') - max_depth: Maximum crawl depth (number, optional, default: 10) - limit: Maximum pages to crawl (number, optional, default: 100) - strategy: Crawl strategy 'all' | 'same-domain' | 'same-hostname' | 'same-origin' (optional, default: 'same-domain') - include_paths: Path patterns to include (string[], optional) - exclude_paths: Path patterns to exclude (string[], optional) - retry: Enable auto-retry on failure (boolean, optional) - poll_seconds: Polling interval for job status (number, optional) - poll_interval_ms: Polling interval in milliseconds (number, optional) - timeout_ms: Job timeout in milliseconds (number, optional) - scrape_options: Nested scrape options for each page (object, optional) Returns: { job_id, status, message } for async jobs Examples: - Recommended: { "url": "https://example.com", "engine": "playwright", "limit": 50 } - Deep crawl: { "url": "https://docs.example.com", "engine": "playwright", "max_depth": 5, "limit": 200 } - Filtered crawl: { "url": "https://blog.example.com", "engine": "playwright", "include_paths": ["/posts/*"], "exclude_paths": ["/admin/*"] }
anycrawl_search
Search the web and optionally scrape results. Best for: Open-ended discovery, finding relevant content. Not recommended for: Known URLs (use anycrawl_scrape) or comprehensive site coverage (use anycrawl_crawl). RECOMMENDED: Use limit=5 for balanced performance and cost. Use 'playwright' engine for scraping results. Usage (parameters): - query: Search query string (string, required) - engine: Search engine 'google' (optional, default: 'google') - limit: Number of results to return (number, optional, default: 5) - offset: Number of results to skip (number, optional, default: 0) - pages: Number of search result pages to process (number, optional) - lang: Language code (string, optional) - country: Country code (string, optional) - safeSearch: Safe search level 0-2 (number, optional) - scrape_options: Options for scraping search results (object, optional) Returns: Array of search results with optional scraped content Examples: - Recommended: { "query": "artificial intelligence news", "limit": 5 } - With scraping: { "query": "TypeScript tutorials", "limit": 5, "scrape_options": { "formats": ["markdown"], "engine": "playwright" } } - Localized search: { "query": "machine learning", "lang": "es", "country": "ES", "limit": 5 }
anycrawl_crawl_status
Get the status of a crawl job. Check progress, completion status, and statistics for an ongoing or completed crawl job. Usage (parameters): - job_id: The crawl job ID (string, required) Returns: { job_id, status, start_time, expires_at, credits_used, total, completed, failed } Examples: - Check status: { "job_id": "crawl_12345" }
anycrawl_crawl_results
Get the results of a completed crawl job. Retrieve the scraped content and metadata from a completed crawl job. Usage (parameters): - job_id: The crawl job ID (string, required) - skip: Number of results to skip (number, optional, default: 0) Returns: { status, total, completed, creditsUsed, next?, data[] } Examples: - Get all results: { "job_id": "crawl_12345" } - Paginated results: { "job_id": "crawl_12345", "skip": 50 }
anycrawl_cancel_crawl
Cancel a running crawl job. Stop an ongoing crawl job and prevent further processing. Usage (parameters): - job_id: The crawl job ID to cancel (string, required) Returns: { success: boolean, message: string } Examples: - Cancel job: { "job_id": "crawl_12345" }
Scrape a single URL and extract content in various formats.
Best for:
- Extracting content from a single page
- Quick data extraction
- Testing specific URLs
Parameters:
- url (required): The URL to scrape
- engine (optional): Scraping engine (auto, playwright, cheerio, puppeteer; default: auto)
- formats (optional): Output formats (markdown, html, text, screenshot, screenshot@fullPage, rawHtml, json)
- proxy (optional): Proxy URL
- timeout (optional): Timeout in milliseconds (default: 300000)
- retry (optional): Whether to retry on failure (default: false)
- wait_for (optional): Wait time for page to load
- include_tags (optional): HTML tags to include
- exclude_tags (optional): HTML tags to exclude
- json_options (optional): Options for JSON extraction
Example:
{
"name": "anycrawl_scrape",
"arguments": {
"url": "https://example.com",
"formats": ["markdown", "html"],
"timeout": 30000
}
}
Start a crawl job to scrape multiple pages from a website. By default this waits for completion and returns aggregated results using the SDK's client.crawl (defaults: poll every 3 seconds, timeout after 60 seconds).
Best for:
- Extracting content from multiple related pages
- Comprehensive website analysis
- Bulk data collection
Parameters:
- url (required): The base URL to crawl
- engine (optional): Scraping engine (auto, playwright, cheerio, puppeteer; default: auto)
- max_depth (optional): Maximum crawl depth (default: 10)
- limit (optional): Maximum number of pages (default: 100)
- strategy (optional): Crawling strategy (all, same-domain, same-hostname, same-origin)
- exclude_paths (optional): URL patterns to exclude
- include_paths (optional): URL patterns to include
- scrape_options (optional): Options for individual page scraping
- poll_seconds (optional): Poll interval seconds for waiting (default: 3)
- timeout_ms (optional): Overall timeout milliseconds for waiting (default: 60000)
Example:
{
"name": "anycrawl_crawl",
"arguments": {
"url": "https://example.com/blog",
"max_depth": 2,
"limit": 50,
"strategy": "same-domain",
"poll_seconds": 3,
"timeout_ms": 60000
}
}
Returns: { "job_id": "...", "status": "completed", "total": N, "completed": N, "creditsUsed": N, "data": [...] }.
Check the status of a crawl job.
Parameters:
- job_id (required): The crawl job ID
Example:
{
"name": "anycrawl_crawl_status",
"arguments": {
"job_id": "7a2e165d-8f81-4be6-9ef7-23222330a396"
}
}
Get results from a crawl job.
Parameters:
- job_id (required): The crawl job ID
- skip (optional): Number of results to skip (for pagination)
Example:
{
"name": "anycrawl_crawl_results",
"arguments": {
"job_id": "7a2e165d-8f81-4be6-9ef7-23222330a396",
"skip": 0
}
}
Cancel a pending crawl job.
Parameters:
- job_id (required): The crawl job ID to cancel
Example:
{
"name": "anycrawl_cancel_crawl",
"arguments": {
"job_id": "7a2e165d-8f81-4be6-9ef7-23222330a396"
}
}
Search the web using AnyCrawl search engine.
Best for:
- Finding specific information across multiple websites
- Research and discovery
- When you don't know which website has the information
Parameters:
- query (required): Search query
- engine (optional): Search engine (google)
- limit (optional): Maximum number of results (default: 10)
- offset (optional): Number of results to skip (default: 0)
- pages (optional): Number of pages to search
- lang (optional): Language code
- country (optional): Country code
- scrape_options (required): Options for scraping search results
- safeSearch (optional): Safe search level (0=off, 1=moderate, 2=strict)
Example:
{
"name": "anycrawl_search",
"arguments": {
"query": "latest AI research papers 2024",
"engine": "google",
"limit": 5,
"scrape_options": {
"formats": ["markdown"]
}
}
}
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"anycrawl": {
"anycrawl-mcp": {
"command": "npx",
"args": [
"-y",
"anycrawl-mcp-server"
],
"env": {
"ANYCRAWL_API_KEY": "<YOUR_TOKEN>",
"ANYCRAWL_BASE_URL": "https://api.anycrawl.dev",
"LOG_LEVEL": "info"
}
}
}
}
}
McpServers
{
"anycrawl-mcp": {
"command": "npx",
"args": [
"-y",
"anycrawl-mcp-server"
],
"env": {
"ANYCRAWL_API_KEY": "<YOUR_TOKEN>",
"ANYCRAWL_BASE_URL": "https://api.anycrawl.dev",
"LOG_LEVEL": "info"
}
}
}
🚀 AnyCrawl MCP Server — Powerful web scraping and crawling for Cursor, Claude, and other LLM clients via the Model Context Protocol (MCP).
Features
- Web Scraping: Extract content from single URLs with multiple output formats
- Website Crawling: Crawl entire websites with configurable depth and limits
- Search Engine Integration: Search the web and optionally scrape results
- Multiple Engines: Auto mode (intelligent selection), Playwright, Cheerio, and Puppeteer
- Flexible Output: Markdown, HTML, text, screenshots, and structured JSON
- Async Operations: Non-blocking crawl jobs with status monitoring
- Error Handling: Robust error handling and logging
- Multiple Modes: STDIO (default), MCP(HTTP), SSE; cloud-ready with Nginx proxy
Installation
Running with npx
ANYCRAWL_API_KEY=YOUR-API-KEY npx -y anycrawl-mcp
Manual installation
npm install -g anycrawl-mcp-server
ANYCRAWL_API_KEY=YOUR-API-KEY anycrawl-mcp
Configuration
AnyCrawl MCP Server supports two deployment modes: Cloud Service (recommended) and Self-Hosted.
Cloud Service (Recommended)
Use the AnyCrawl cloud service at mcp.anycrawl.dev. No server setup required.
```bash
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



