Webfetch Mcp

by simonediroma

374 downloads Not rated yet
GitHub

About

# webfetch-mcp ![Python](https://img.shields.io/badge/python-3.10%2B-blue) ![License](https://img.shields.io/badge/license-MIT-green) ![MCP](https://img.shields.io/badge/MCP-compatible-purple) [![simonediroma/webfetch_mcp MCP…

Explore

| Feature | Description |
|---------|-------------|
| Domain-scoped headers | Different auth headers per domain; global * fallback |
| Per-call headers | The client (or you) can inject extra headers for a single request |
| YAML config | Single readable file controls headers, timeouts, retries, proxies, and output formats |
| Configurable timeout | Per-domain request timeout (default 30 s) |
| Retry with backoff | Auto-retry on HTTP 5xx or network errors, with exponential backoff |
| Per-domain proxy | Route traffic through a different proxy per domain |
| Output formats | raw, markdown, trafilatura (main content), json (pretty-print), lighthtml (minimal HTML) |
| JSON auto-detection | Responses with application/json Content-Type are pretty-printed automatically |
| Metadata extraction | Extracts title, author, date, source via trafilatura (opt-in per domain) |
| Bot-block detection | Detects Cloudflare / CAPTCHA blocks; optionally retries with a Chrome User-Agent |
| Prompt-injection sanitization | Scans fetched content for injection patterns; flag or strip mode |
| CSS selector extraction | Extract specific HTML elements before format conversion, configurable per domain or per call |
| Redirect tracing | Optionally record and display the full redirect chain in the summary |
| Response assertions | assert_status / assert_contains raise an error on mismatch — useful for CI/CD smoke tests |
| Header injection protection | Validates headers for control characters (\r, \n, NUL) |
| Response truncation | max_bytes cap to avoid filling the assistant's context window |
| Detailed response summary | Every response includes a structured summary (status, elapsed ms, injected headers, format, etc.) |
| JS rendering (Playwright) | Render JavaScript-heavy SPAs with headless Chromium before extracting content; configurable globally, per-domain, or per-call |
| lighthtml output format | Strips <style>, <script> (except JSON-LD), comments, and all tag attributes — returns minimal bare HTML structure |

---

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Webfetch Mcp
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

- Python 3.10+
- Any MCP-compatible AI assistant (Claude Code, Cursor, Continue, Zed, etc.)

---

Copy the example and edit it:

cp webfetch.yaml.example webfetch.yaml

Point the server at it:


Copy .env.example and fill in your values:

bash
cp .env.example .env

WEBFETCH_HEADERS — domain-scoped request headers (single-line JSON):

env
WEBFETCH_HEADERS={"": {"User-Agent": "MyBot/1.0"}, "example.com": {"X-Auth-Token": "your-token"}}

WEBFETCH_OUTPUT — domain-scoped output format (single-line JSON):

env
WEBFETCH_OUTPUT={"
": "raw", "example.com": "trafilatura", "news.com": "markdown"}

WEBFETCH_SELECTORS — domain-scoped CSS selector (single-line JSON):

env
WEBFETCH_SELECTORS={"example.com": "article.main-content", "news.com": "div#article-body"}

> When WEBFETCH_CONFIG is set, the env vars above are ignored entirely.

---

bash

Parameter

Type

url

`str`

method

`str`

body

`str \

extra_headers

`dict \

extract_text

`bool`

max_bytes

`int`

follow_redirects

`bool`

output_format

`str \

css_selector

`str \

trace_redirects

`bool`

assert_status

`int \

assert_contains

`str \

render_js

`bool \

Most AI assistants expose both their built-in WebFetch and any registered MCP tools. To ensure mcp__webfetch__fetch is always preferred:

All parameters are optional except url.

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| url | str | — | URL to fetch |
| method | str | "GET" | HTTP verb (GET, POST, PUT, DELETE, …) |
| body | str \| None | None | Request body for POST/PUT |
| extra_headers | dict \| None | None | Per-call headers merged on top of domain headers |
| extract_text | bool | False | Strip HTML tags, return plain text (legacy; overrides output_format) |
| max_bytes | int | 0 | Truncate response to N characters (0 = unlimited) |
| follow_redirects | bool | True | Follow HTTP redirects |
| output_format | str \| None | None | Per-call format override: "raw", "markdown", "trafilatura", "json", "lighthtml" |
| css_selector | str \| None | None | CSS selector to extract HTML element(s) before format conversion (e.g. "article", "#main") |
| trace_redirects | bool | False | Display the full redirect chain in the summary |
| assert_status | int \| None | None | Raise an error if the response status code does not match this value |
| assert_contains | str \| None | None | Raise an error if this string is not found in the response body (case-sensitive) |
| render_js | bool \| None | None | Render the page with headless Chromium (executes JS, waits for network idle). Requires playwright. |

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "webfetch mcp": {
            "webfetch": {
                "command": "/absolute/path/to/.venv/bin/python",
                "args": [
                    "/absolute/path/to/server.py"
                ],
                "env": {
                    "WEBFETCH_CONFIG": "/absolute/path/to/webfetch.yaml"
                }
            }
        }
    }
}

McpServers

{
    "webfetch": {
        "command": "/absolute/path/to/.venv/bin/python",
        "args": [
            "/absolute/path/to/server.py"
        ],
        "env": {
            "WEBFETCH_CONFIG": "/absolute/path/to/webfetch.yaml"
        }
    }
}

Python
License
MCP
simonediroma/webfetch_mcp MCP server

A local Python MCP server that replaces your AI assistant's built-in WebFetch tool with a fully configurable HTTP client — supporting domain-scoped headers, retries, proxies, timeouts, output formats, bot-block detection, and prompt-injection sanitization, all without touching a single line of your assistant's config beyond registering the server.

Why

The built-in WebFetch tool available in most AI assistants (Claude Code, Cursor, Continue, Zed, etc.) sends requests without custom headers, which means it gets blocked by bot-protection systems (Akamai, Cloudflare, paywalls, etc.) and can't authenticate against APIs that require domain-specific tokens.

This server is a drop-in replacement: it exposes the same fetch tool to any MCP-compatible AI assistant, but enriches every outbound request with the right headers, format, and retry strategy based on the target domain — automatically, without you having to configure headers every time.

---

Features

| Feature | Description |
|---------|-------------|
| Domain-scoped headers | Different auth headers per domain; global fallback |
| Per-call headers | The client (or you) can inject extra headers for a single request |
| YAML config | Single readable file controls headers, timeouts, retries, proxies, and output formats |
| Configurable timeout | Per-domain request timeout (default 30 s) |
| Retry with backoff | Auto-retry on HTTP 5xx or network errors, with exponential backoff |
| Per-domain proxy | Route traffic through a different proxy per domain |
| Output formats | raw, markdown, trafilatura (main content), json (pretty-print), lighthtml (minimal HTML) |
| JSON auto-detection | Responses with application/json Content-Type are pretty-printed automatically |
| Metadata extraction | Extracts title, author, date, source via trafilatura (opt-in per domain) |
| Bot-block detection | Detects Cloudflare / CAPTCHA blocks; optionally retries with a Chrome User-Agent |
| Prompt-injection sanitization | Scans fetched content for injection patterns; flag or strip mode |
| CSS selector extraction | Extract specific HTML elements before format conversion, configurable per domain or per call |
| Redirect tracing | Optionally record and display the full redirect chain in the summary |
| Response assertions | assert_status / assert_contains raise an error on mismatch — useful for CI/CD smoke tests |
| Header injection protection | Validates headers for control characters (\r, \n, NUL) |
| Response truncation | max_bytes cap to avoid filling the assistant's context window |
| Detailed response summary | Every response includes a structured summary (status, elapsed ms, injected headers, format, etc.) |
| JS rendering (Playwright) | Render JavaScript-heavy SPAs with headless Chromium before extracting content; configurable globally, per-domain, or per-call |
| lighthtml output format | Strips <style>, <script> (except JSON-LD), comments, and all tag attributes — returns minimal bare HTML structure |

---

Requirements

- Python 3.10+
- Any MCP-compatible AI assistant (Claude Code, Cursor, Continue, Zed, etc.)

---

Quick start

git clone https://github.com/simonediroma/webfetch_mcp.git
cd webfetch_mcp

Mac / Linux

python -m venv .venv && .venv/bin/pip install -r requirements.txt

Windows

python -m venv .venv && .venv\Scripts\pip install -r requirements.txt

cp webfetch.yaml.example webfetch.yaml # then edit with your tokens

Then register the server in your AI assistant config and restart. Done.

---

Installation

git clone https://github.com/simonediroma/webfetch_mcp.git
cd webfetch_mcp

python -m venv .venv

Windows

.venv\Scripts\pip install -r requirements.txt

Mac / Linux

.venv/bin/pip install -r requirements.txt

requirements.txt installs:

mcp[cli]>=1.0.0
httpx>=0.27.0
python-dotenv>=1.0.0
markdownify>=0.12.0
trafilatura>=1.12.0
pyyaml>=6.0
beautifulsoup4>=4.12.0

Optional — JS rendering requires Playwright:

pip install playwright && playwright install chromium

---

Configuration

There are two ways to configure the server. YAML is recommended — it supports all options. The legacy environment variable approach still works for simple cases.

Option A — YAML config file (recommended)

Copy the example and edit it:

cp webfetch.yaml.example webfetch.yaml

Point the server at it:

# In your shell profile, or in the MCP server env block (see Registration below)
export WEBFETCH_CONFIG=/absolute/path/to/webfetch.yaml
Full YAML reference
# Global defaults — applied to every request unless overridden
global:
  headers:
    User-Agent: "MyBot/1.0"
  output_format: raw       # raw | markdown | trafilatura | json | lighthtml
  timeout: 30              # seconds
  retry:
    attempts: 1            # 1 = no retry
    backoff: 2.0           # exponential multiplier (1s → 2s → 4s …)
  proxy: null              # e.g. "http://proxy.corp:8080"
  extract_metadata: false  # true = prepend title/author/date to content
  sanitize_content: false  # false | "flag" | "strip"
  bot_block_detection: false  # false | "report" | "retry"
  css_selector: null       # CSS selector to extract element(s) before format conversion
  render_js: false           # true = render JS via headless Chromium (requires playwright)

Per-domain overrides — only the fields you list are overridden

domains: example.com: headers: X-Akamai-Token: "your-token-here" output_format: trafilatura timeout: 60 retry: attempts: 3 backoff: 2.0

news-site.com:
output_format: markdown
bot_block_detection: retry # auto-retry with Chrome UA if blocked
css_selector: "article.main-content" # extract only the article body

internal.corp:
proxy: "http://proxy.corp:8080"
headers:
Authorization: "Bearer my-internal-token"

api.example.com:
output_format: json
timeout: 10
retry:
attempts: 5
backoff: 1.5

Domain matching uses suffix rules: example.com matches both example.com and www.example.com. When multiple domains match, the most specific (longest) key wins. Global settings are always applied first, then overridden by increasingly specific domain rules.

---

Option B — Environment variables (legacy)

Copy .env.example and fill in your values:

cp .env.example .env

WEBFETCH_HEADERS — domain-scoped request headers (single-line JSON):

WEBFETCH_HEADERS={"": {"User-Agent": "MyBot/1.0"}, "example.com": {"X-Auth-Token": "your-token"}}

WEBFETCH_OUTPUT — domain-scoped output format (single-line JSON):

WEBFETCH_OUTPUT={"*": "raw", "example.com": "trafilatura", "news.com": "markdown"}

WEBFETCH_SELECTORS — domain-scoped CSS selector (single-line JSON):

WEBFETCH_SELECTORS={"example.com": "article.main-content", "news.com": "div#article-body"}

> When WEBFETCH_CONFIG is set, the env vars above are ignored entirely.

---

Registering with your AI assistant

Most AI assistants use a mcpServers block in a JSON settings file. The format is the same across assistants — only the file location differs.

Claude Code

Add to ~/.claude/settings.json:

…

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.