just-every/mcp-read-website-fast
About
Fast, token-efficient web content extraction that converts websites to clean Markdown. Features Mozilla Readability, smart caching, polite crawling with robots.txt support, and concurrent fetching with minimal dependencies.
Details
- Author
- just-every
- Categories
- Web Scraping, Community, Other
Jump to
Setup
Install just-every/mcp-read-website-fast in your MCP client (Claude Desktop, Cursor, Windsurf, and others).
Repository: https://github.com/just-every/mcp-read-website-fast
Follow the installation instructions in the repository README, then restart your MCP client.
Fast, token-efficient web content extraction for AI agents - converts websites to clean Markdown.
Existing MCP web crawlers are slow and consume large quantities of tokens. This pauses the development process and provides incomplete results as LLMs need to parse whole web pages.
This MCP package fetches web pages locally, strips noise, and converts content to clean Markdown while preserving links. Designed for Claude Code, IDEs and LLM pipelines with minimal token footprint. Crawl sites locally with minimal dependencies.
Note:This package now uses@just-every/crawlfor its core crawling and markdown conversion functionality.
- Fast startupusing official MCP SDK with lazy loading for optimal performance
- Content extractionusing Mozilla Readability (same as Firefox Reader View)
- HTML to Markdownconversion with Turndown + GFM support
- Smart cachingwith SHA-256 hashed URLs
- Polite crawlingwith robots.txt support and rate limiting
- Concurrent fetchingwith configurable depth crawling
- Stream-first designfor low memory usage
- Link preservationfor knowledge graphs
- Optional chunkingfor downstream processing
claude mcp add read-website-fast -s user -- npx -y @just-every/mcp-read-website-fast
code --add-mcp '{"name":"read-website-fast","command":"npx","args":["-y","@just-every/mcp-read-website-fast"]}'
cursor://anysphere.cursor-deeplink/mcp/install?name=read-website-fast&config=eyJyZWFkLXdlYnNpdGUtZmFzdCI6eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIkBqdXN0LWV2ZXJ5L21jcC1yZWFkLXdlYnNpdGUtZmFzdCJdfX0=
Settings → Tools → AI Assistant → Model Context Protocol (MCP) → Add
{"command":"npx","args":["-y","@just-every/mcp-read-website-fast"]}
Or, in the chat window, type /add and fill in the same JSON—both paths land the server in a single step. 
{ "mcpServers": { "read-website-fast": { "command": "npx", "args": ["-y", "@just-every/mcp-read-website-fast"] } } }
Drop this into your client’s mcp.json (e.g. .vscode/mcp.json, ~/.cursor/mcp.json, or .mcp.json for Claude).
- Fast startupusing official MCP SDK with lazy loading for optimal performance
- Content extractionusing Mozilla Readability (same as Firefox Reader View)
- HTML to Markdownconversion with Turndown + GFM support
- Smart cachingwith SHA-256 hashed URLs
- Polite crawlingwith robots.txt support and rate limiting
- Concurrent fetchingwith configurable depth crawling
- Stream-first designfor low memory usage
- Link preservationfor knowledge graphs
- Optional chunkingfor downstream processing
- read_website- Fetches a webpage and converts it to clean markdown
- Parameters:
- url(required): The HTTP/HTTPS URL to fetch
- pages(optional): Maximum number of pages to crawl (default: 1, max: 100)
- read-website-fast://status- Get cache statistics
- read-website-fast://clear-cache- Clear the cache directory
npm run dev fetch https://example.com/article
npm run dev fetch https://example.com --depth 2 --concurrency 5
# Markdown only (default) npm run dev fetch https://example.com # JSON output with metadata npm run dev fetch https://example.com --output json # Both URL and markdown npm run dev fetch https://example.com --output both
- -p, --pages <number>- Maximum number of pages to crawl (default: 1)
- -c, --concurrency <number>- Max concurrent requests (default: 3)
- --no-robots- Ignore robots.txt
- --all-origins- Allow cross-origin crawling
- -u, --user-agent <string>- Custom user agent
- --cache-dir <path>- Cache directory (default: .cache)
- -t, --timeout <ms>- Request timeout in milliseconds (default: 30000)
- -o, --output <format>- Output format: json, markdown, or both (default: markdown)
The MCP server includes automatic restart capability by default for improved reliability:
- Automatically restarts the server if it crashes
- Handles unhandled exceptions and promise rejections
- Implements exponential backoff (max 10 attempts in 1 minute)
- Logs all restart attempts for monitoring
- Gracefully handles shutdown signals (SIGINT, SIGTERM)
For development/debugging without auto-restart:
# Run directly without restart wrapper npm run serve:dev
mcp/ ├── src/ │ ├── crawler/ # URL fetching, queue management, robots.txt │ ├── parser/ # DOM parsing, Readability, Turndown conversion │ ├── cache/ # Disk-based caching with SHA-256 keys │ ├── utils/ # Logger, chunker utilities │ ├── index.ts # CLI entry point │ ├── serve.ts # MCP server entry point │ └── serve-restart.ts # Auto-restart wrapper
# Run in development mode npm run dev fetch https://example.com # Build for production npm run build # Run tests npm test # Type checking npm run typecheck # Linting npm run lint
- Fork the repository
- Create a feature branch
- Add tests for new functionality
- Submit a pull request
- Increase timeout with-tflag
- Check network connectivity
- Verify URL is accessible
- Some sites block automated access
- Try custom user agent with-uflag
- Check if site requires JavaScript (not supported)
Browser MCP server for AI agents to automate web pages with Puppeteer, accessibility-tree actions, optional vision mode, and cross-platform browser control.
A MCP server to retrieve up-to-date jobs from company career sites.
Fetch the content of a remote URL as Markdown with Jina Reader.
High-quality screenshot capture optimized for Claude Vision API. Automatically tiles full pages into 1072x1072 chunks (1.15 megapixels) with configurable viewports and wait strategies for dynamic content.
An MCP server using Playwright for browser automation and webscrapping
Secure fetch to prevent access to local resources
MCP Server to let Claude / your AI control the browser
A MCP server that provides comprehensive website snapshot capabilities using Playwright. This server enables LLMs to capture and analyze web pages through structured accessibility snapshots, network monitoring, and console message collection.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.




