ArXiv

Recommended

by blazickjp

879 555 downloads Not rated yet Apache-2.0

About

Search for papers and access their content in a programmatic way.

Details

Repository
blazickjp/arxiv-mcp-server
License
Apache-2.0

Explore

- Search arXiv with filters for categories, date ranges, and boolean operators.
- Download paper content (HTML first, with PDF fallback).
- Read paper content with paging support for large documents.
- List all locally downloaded papers.
- Local storage for faster repeated access.
- Pre‑built research prompts for paper analysis.

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name ArXiv
    Command (node, npx, python, etc.) uvx
    Arguments
    • Argument 1 arxiv-mcp-server

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

uvxreuses cached tool environments. Force it to resolve the current PyPI release with a supported interpreter, then restart your MCP client:

uvx --python 3.11 --refresh-package arxiv-mcp-server arxiv-mcp-server

If your client still launches an older environment, add"--python", "3.11"before"arxiv-mcp-server"in itsargsarray.

Desktop applications do not always inherit the samePATHas your terminal. Ifuvx arxiv-mcp-serverworks in a terminal but the client reports that the server failed to connect, find the executable's absolute path:

# Windows PowerShell (Get-Command uvx).Source

Replace"command": "uvx"with the returned absolute path, then restart the client. Keep theargsvalue unchanged.

To placearxiv-mcp-serveron yourPATHinstead of launching it throughuvx:

If the command is not immediately available, runuv tool update-shelland restart the terminal. Afterward, use"command": "arxiv-mcp-server"and omit the package name fromargs.

The repository now packages the same MCP server and research skill for both major plugin systems:

Direct MCP installation is the shortest path. Install the plugin when you also want the research workflow that steers the client toward focused searches, bounded reads, citation traversal, and section-level LaTeX retrieval.

Ask your MCP client to callsearch_paperswith:

{ "query": "\"Kolmogorov-Arnold Networks\"", "categories": ["cs.LG", "cs.AI"], "max_results": 5, "sort_by": "date" }
{ "paper_id": "2404.19756" }
{ "paper_id": "2404.19756", "max_chars": 12000 }

Then page through the cached content withread_paper:

{ "paper_id": "2404.19756", "start": 0, "max_chars": 12000 }

Large-content responses includecontent_length,returned_chars,next_start, andis_truncated. Passnext_startinto the next call to continue reading.

{ "paper_id": "1706.03762" }

Get the first page of its section outline withlist_paper_latex_sections:

{ "paper_id": "1706.03762", "start": 0, "max_sections": 100 }

Then callget_paper_latex_sectionusing an ID from that outline:

{ "paper_id": "1706.03762", "section_id": "3.2", "max_chars": 12000 }

LaTeX archives are validated, size-limited, and cached locally before content is returned.

Choose the install variant that matches the features you need:

# Base server uv tool install arxiv-mcp-server # Base server plus PDF conversion uv tool install "arxiv-mcp-server[pdf]" # Base server plus local semantic search uv tool install "arxiv-mcp-server[pro]"

If the base tool is already installed, reinstall the selected variant:

uv tool install --force "arxiv-mcp-server[pdf]"

Thepdfextra installspymupdf4llmandpymupdf-layoutfor papers without usable arXiv HTML. Theproextra adds local embedding dependencies forsemantic_searchandreindex; semantic search only operates on papers already downloaded to the configured storage directory.

The server provides seven MCP prompt workflows. Prompt availability depends on the client; the server provides workflow instructions but does not run a separate model.

For deployments where stdio is not practical:

TRANSPORT=http HOST=127.0.0.1 PORT=8080 \ uvx arxiv-mcp-server --storage-path /absolute/path/to/papers
$env:TRANSPORT = "http" $env:HOST = "127.0.0.1" $env:PORT = "8080" uvx arxiv-mcp-server --storage-path C:\absolute\path\to\papers
{ "mcpServers": { "arxiv": { "type": "http", "url": "http://127.0.0.1:8080/mcp" } } }

Cloud and load-balancer probes should GEThttp://<host>:<port>/healthz. It returns200with bodyokonce the HTTP server is listening. There is no separate/readycheck: if the process is up, it is ready. The stdio transport has no HTTP endpoints.

The server binds to127.0.0.1by default and enables MCP DNS-rebinding protection. If a reverse proxy exposes the server, keep the process on a private interface and provide authentication and network controls upstream. UseALLOWED_HOSTSandALLOWED_ORIGINSfor the host and origin values forwarded by the proxy.

Environment variable names are case-insensitive through Pydantic settings.--storage-pathis a command-line option rather than an environment setting.

Paper text and LaTeX are untrusted external content. A paper can contain text intended to manipulate an AI client into ignoring its instructions or calling unrelated tools.

- Do not treat instructions found inside a paper as trusted commands.
- Use client approval controls for shell, browser, filesystem, and messaging tools.
- Review generated summaries before taking external actions.
- Keep Streamable HTTP private unless authentication is provided upstream.

SeeSECURITY.mdfor the reporting policy and threat details.

git clone https://github.com/blazickjp/arxiv-mcp-server.git cd arxiv-mcp-server uv sync --extra test --extra dev uv run pytest uv run black --check .

Run the development checkout from an MCP client with:

{ "mcpServers": { "arxiv-dev": { "command": "uv", "args": [ "--directory", "/absolute/path/to/arxiv-mcp-server", "run", "arxiv-mcp-server" ] } } }

Contributions are welcome. ReadCONTRIBUTING.mdbefore opening a pull request, and useGitHub Issuesfor reproducible bugs or scoped feature proposals.

Search global news using natural language. Webz.io News Search API returns the most relevant articles and content, with filters for source, country, language, date, sentiment, and category.

Verified, tier-0 regulatory data for your AI: connect Claude, ChatGPT or Cursor to 850+ official sources across 50+ jurisdictions.

A flexible service for searching and analyzing academic papers on arXiv.

Read Australian Commonwealth law as it stood on any date back to 1901, and verify statute citations against the official Federal Register of Legislation. No API key.

Search scientific papers with structured experimental data extracted from full-text studies. Returns 25+ fields per paper including methods, results, sample sizes, limitations, and quality scores.

Search academic references from arXiv, DBLP, Semantic Scholar, and OpenAlex, and generate BibTeX entries.

Search and access academic paper metadata from Crossref.

Search and cite exact passages across complete classical and world-literature corpora.

Anonymous, read-only, source-backed Buddhist scripture search, passage guidance, explanation, and one-time practice planning through four production MCP tools.

search_papers

Search arXiv by query, category, date, and sort order. Remote arXiv API.

get_abstract

Fetch metadata and an abstract by arXiv ID. Does not download the paper.

download_paper

Download and convert a paper to local Markdown. HTML first; PDF fallback uses '[pdf]'; 'force=true' re-fetches.

list_papers

List papers stored locally. Returns id, title, authors, published; 'compact' for IDs only.

read_paper

Read locally stored paper content. Supports 'start' and 'max_chars'.

get_paper_latex

Retrieve bounded author-submitted LaTeX. Remote arXiv source archive.

list_paper_latex_sections

Return a paginated LaTeX outline. Supports 'start' and 'max_sections'.

get_paper_latex_section

Read one bounded LaTeX section. Select by outline ID or exact title.

citation_graph

Fetch references and citing papers. Remote Semantic Scholar API; optional 'SEMANTIC_SCHOLAR_API_KEY'.

export_citations

Export BibTeX for one or more arXiv IDs. Authoritative arXiv metadata.

watch_topic

Save or update an arXiv topic watch. Stored locally.

list_watches

List saved topic watches. Read-only; does not advance last_checked.

check_alerts

Check saved watches for new papers. Returns papers since the last check.

unwatch_topic

Delete a saved topic watch. Exact topic match; not-found if missing.

semantic_search

Search downloaded papers by semantic similarity. Requires '[pro]'.

reindex

Rebuild the local semantic index. Requires '[pro]'.

uvxreuses cached tool environments. Force it to resolve the current PyPI release with a supported interpreter, then restart your MCP client:

uvx --python 3.11 --refresh-package arxiv-mcp-server arxiv-mcp-server

If your client still launches an older environment, add"--python", "3.11"before"arxiv-mcp-server"in itsargsarray.

Desktop applications do not always inherit the samePATHas your terminal. Ifuvx arxiv-mcp-serverworks in a terminal but the client reports that the server failed to connect, find the executable's absolute path:

# Windows PowerShell (Get-Command uvx).Source

Replace"command": "uvx"with the returned absolute path, then restart the client. Keep theargsvalue unchanged.

To placearxiv-mcp-serveron yourPATHinstead of launching it throughuvx:

If the command is not immediately available, runuv tool update-shelland restart the terminal. Afterward, use"command": "arxiv-mcp-server"and omit the package name fromargs.

The repository now packages the same MCP server and research skill for both major plugin systems:

Direct MCP installation is the shortest path. Install the plugin when you also want the research workflow that steers the client toward focused searches, bounded reads, citation traversal, and section-level LaTeX retrieval.

Ask your MCP client to callsearch_paperswith:

{ "query": "\"Kolmogorov-Arnold Networks\"", "categories": ["cs.LG", "cs.AI"], "max_results": 5, "sort_by": "date" }
{ "paper_id": "2404.19756" }
{ "paper_id": "2404.19756", "max_chars": 12000 }

Then page through the cached content withread_paper:

{ "paper_id": "2404.19756", "start": 0, "max_chars": 12000 }

Large-content responses includecontent_length,returned_chars,next_start, andis_truncated. Passnext_startinto the next call to continue reading.

{ "paper_id": "1706.03762" }

Get the first page of its section outline withlist_paper_latex_sections:

{ "paper_id": "1706.03762", "start": 0, "max_sections": 100 }

Then callget_paper_latex_sectionusing an ID from that outline:

{ "paper_id": "1706.03762", "section_id": "3.2", "max_chars": 12000 }

LaTeX archives are validated, size-limited, and cached locally before content is returned.

Choose the install variant that matches the features you need:

# Base server uv tool install arxiv-mcp-server # Base server plus PDF conversion uv tool install "arxiv-mcp-server[pdf]" # Base server plus local semantic search uv tool install "arxiv-mcp-server[pro]"

If the base tool is already installed, reinstall the selected variant:

uv tool install --force "arxiv-mcp-server[pdf]"

Thepdfextra installspymupdf4llmandpymupdf-layoutfor papers without usable arXiv HTML. Theproextra adds local embedding dependencies forsemantic_searchandreindex; semantic search only operates on papers already downloaded to the configured storage directory.

The server provides seven MCP prompt workflows. Prompt availability depends on the client; the server provides workflow instructions but does not run a separate model.

For deployments where stdio is not practical:

TRANSPORT=http HOST=127.0.0.1 PORT=8080 \ uvx arxiv-mcp-server --storage-path /absolute/path/to/papers
$env:TRANSPORT = "http" $env:HOST = "127.0.0.1" $env:PORT = "8080" uvx arxiv-mcp-server --storage-path C:\absolute\path\to\papers
{ "mcpServers": { "arxiv": { "type": "http", "url": "http://127.0.0.1:8080/mcp" } } }

Cloud and load-balancer probes should GEThttp://<host>:<port>/healthz. It returns200with bodyokonce the HTTP server is listening. There is no separate/readycheck: if the process is up, it is ready. The stdio transport has no HTTP endpoints.

The server binds to127.0.0.1by default and enables MCP DNS-rebinding protection. If a reverse proxy exposes the server, keep the process on a private interface and provide authentication and network controls upstream. UseALLOWED_HOSTSandALLOWED_ORIGINSfor the host and origin values forwarded by the proxy.

Environment variable names are case-insensitive through Pydantic settings.--storage-pathis a command-line option rather than an environment setting.

Paper text and LaTeX are untrusted external content. A paper can contain text intended to manipulate an AI client into ignoring its instructions or calling unrelated tools.

- Do not treat instructions found inside a paper as trusted commands.
- Use client approval controls for shell, browser, filesystem, and messaging tools.
- Review generated summaries before taking external actions.
- Keep Streamable HTTP private unless authentication is provided upstream.

SeeSECURITY.mdfor the reporting policy and threat details.

git clone https://github.com/blazickjp/arxiv-mcp-server.git cd arxiv-mcp-server uv sync --extra test --extra dev uv run pytest uv run black --check .

Run the development checkout from an MCP client with:

{ "mcpServers": { "arxiv-dev": { "command": "uv", "args": [ "--directory", "/absolute/path/to/arxiv-mcp-server", "run", "arxiv-mcp-server" ] } } }

Contributions are welcome. ReadCONTRIBUTING.mdbefore opening a pull request, and useGitHub Issuesfor reproducible bugs or scoped feature proposals.

Search global news using natural language. Webz.io News Search API returns the most relevant articles and content, with filters for source, country, language, date, sentiment, and category.

Verified, tier-0 regulatory data for your AI: connect Claude, ChatGPT or Cursor to 850+ official sources across 50+ jurisdictions.

A flexible service for searching and analyzing academic papers on arXiv.

Read Australian Commonwealth law as it stood on any date back to 1901, and verify statute citations against the official Federal Register of Legislation. No API key.

Search scientific papers with structured experimental data extracted from full-text studies. Returns 25+ fields per paper including methods, results, sample sizes, limitations, and quality scores.

Search academic references from arXiv, DBLP, Semantic Scholar, and OpenAlex, and generate BibTeX entries.

Search and access academic paper metadata from Crossref.

Search and cite exact passages across complete classical and world-literature corpora.

Anonymous, read-only, source-backed Buddhist scripture search, passage guidance, explanation, and one-time practice planning through four production MCP tools.

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "arxiv": {
            "cwd": "string",
            "env": {},
            "args": [
                "arxiv-mcp-server"
            ],
            "shell": false,
            "command": "uvx"
        }
    }
}

Linux

{
    "cwd": "string",
    "env": [],
    "args": [
        "arxiv-mcp-server"
    ],
    "shell": false,
    "command": "uvx"
}

Macos

{
    "cwd": "string",
    "env": [],
    "args": [
        "arxiv-mcp-server"
    ],
    "shell": false,
    "command": "uvx"
}

Windows

{
    "cwd": "string",
    "env": [],
    "args": [
        "/c",
        "uvx",
        "arxiv-mcp-server"
    ],
    "shell": false,
    "command": "cmd"
}

An MCP server for searching arXiv, downloading papers, reading bounded full text, retrieving original LaTeX by section, following citation graphs, and maintaining research alerts.

It runs locally over stdio by default. Papers and indexes stay on your machine; search, source retrieval, citation graphs, and downloads call their respective external services.

The command-based integrations requireuv, which providesuvx. Choose your client below; no repository clone or Python environment setup is required.

claude mcp add --transport stdio --scope user arxiv -- uvx arxiv-mcp-server

For the richer plugin integration—which installs the MCP connection plus the bundled arXiv research skill—register this repository as a marketplace and install the plugin:

claude plugin marketplace add blazickjp/arxiv-mcp-server claude plugin install arxiv-mcp-server@arxiv-mcp

Verify the direct MCP installation withclaude mcp get arxiv. Restart Claude Code or run/reload-pluginsafter installing the plugin.

codex mcp add arxiv -- uvx arxiv-mcp-server

Or install the MCP connection and bundled research skill as a Codex plugin:

codex plugin marketplace add blazickjp/arxiv-mcp-server codex plugin add arxiv-mcp-server@arxiv-mcp

Verify the direct MCP installation withcodex mcp get arxiv. Codex CLI, the Codex IDE extension, and Codex in the ChatGPT desktop app share this MCP configuration.

Add the server, approve the discovered tools, and test the saved connection:

hermes mcp add arxiv --command uvx --args arxiv-mcp-server hermes mcp test arxiv

Use theAdd to Kiro,Install in VS Code, orInstall in VS Code Insidersbutton above.

For the richer Kiro Power integration, open thePowerspanel, chooseAdd Custom Power → Import power from GitHub, and enter:

https://github.com/blazickjp/arxiv-mcp-server

The Power installs the MCP connection frommcp.jsonand adds focused arXiv research guidance. Kiro users who prefer manual configuration can place the generic configuration below in.kiro/settings/mcp.jsonfor one workspace or~/.kiro/settings/mcp.jsonfor all workspaces.

macOS users can install a bundled.mcpbextension from thelatest GitHub release:

- Apple Silicon:arxiv-mcp-server-darwin-arm64-<version>.mcpb
- Intel:arxiv-mcp-server-darwin-x86_64-<version>.mcpb

Double-click the bundle, drag it into Claude Desktop, or openSettings → Extensions → Advanced settings → Install Extension…. The bundle includes the server dependencies and requires CPython 3.11.x.

Add this stdio configuration to clients that accept themcpServersJSON shape, such as Claude Desktop and Kiro. Other clients may use a top-levelserversobject, TOML, or their own settings UI; consult the client's MCP documentation.

{ "mcpServers": { "arxiv": { "type": "stdio", "command": "uvx", "args": ["arxiv-mcp-server"] } } }

The default paper directory is~/.arxiv-mcp-server/papers. To choose another directory, append"--storage-path", "/absolute/path/to/papers"toargs.

For older papers that require PDF conversion, run the package with its PDF extra:

{ "mcpServers": { "arxiv": { "type": "stdio", "command": "uvx", "args": [ "--from", "arxiv-mcp-server[pdf]", "arxiv-mcp-server" ] } } }

The supported package is published on PyPI. An unrelated npm package uses the same name, so do not install this server with npm, pnpm, ornpx arxiv-mcp-server.

If an existing installation is missing newer tools

uvxreuses cached tool environments. Force it to resolve the current PyPI release with a supported interpreter, then restart your MCP client:

uvx --python 3.11 --refresh-package arxiv-mcp-server arxiv-mcp-server

If your client still launches an older environment, add"--python", "3.11"before"arxiv-mcp-server"in itsargsarray.

Desktop applications do not always inherit the samePATHas your terminal. Ifuvx arxiv-mcp-serverworks in a terminal but the client reports that the server failed to connect, find the executable's absolute path:

# Windows PowerShell (Get-Command uvx).Source

Replace"command": "uvx"with the returned absolute path, then restart the client. Keep theargsvalue unchanged.

To placearxiv-mcp-serveron yourPATHinstead of launching it throughuvx:

If the command is not immediately available, runuv tool update-shelland restart the terminal. Afterward, use"command": "arxiv-mcp-server"and omit the package name fromargs.

The repository now packages the same MCP server and research skill for both major plugin systems:

Direct MCP installation is the shortest path. Install the plugin when you also want the research workflow that steers the client toward focused searches, bounded reads, citation traversal, and section-level LaTeX retrieval.

Ask your MCP client to callsearch_paperswith:

{ "query": "\"Kolmogorov-Arnold Networks\"", "categories": ["cs.LG", "cs.AI"], "max_results": 5, "sort_by": "date" }
{ "paper_id": "2404.19756" }
{ "paper_id": "2404.19756", "max_chars": 12000 }

Then page through the cached content withread_paper:

{ "paper_id": "2404.19756", "start": 0, "max_chars": 12000 }

Large-content responses includecontent_length,returned_chars,next_start, andis_truncated. Passnext_startinto the next call to continue reading.

{ "paper_id": "1706.03762" }

Get the first page of its section outline withlist_paper_latex_sections:

{ "paper_id": "1706.03762", "start": 0, "max_sections": 100 }

Then callget_paper_latex_sectionusing an ID from that outline:

{ "paper_id": "1706.03762", "section_id": "3.2", "max_chars": 12000 }

LaTeX archives are validated, size-limited, and cached locally before content is returned.

Choose the install variant that matches the features you need:

# Base server uv tool install arxiv-mcp-server # Base server plus PDF conversion uv tool install "arxiv-mcp-server[pdf]" # Base server plus local semantic search uv tool install "arxiv-mcp-server[pro]"

If the base tool is already installed, reinstall the selected variant:

uv tool install --force "arxiv-mcp-server[pdf]"

Thepdfextra installspymupdf4llmandpymupdf-layoutfor papers without usable arXiv HTML. Theproextra adds local embedding dependencies forsemantic_searchandreindex; semantic search only operates on papers already downloaded to the configured storage directory.

The server provides seven MCP prompt workflows. Prompt availability depends on the client; the server provides workflow instructions but does not run a separate model.

For deployments where stdio is not practical:

TRANSPORT=http HOST=127.0.0.1 PORT=8080 \ uvx arxiv-mcp-server --storage-path /absolute/path/to/papers
$env:TRANSPORT = "http" $env:HOST = "127.0.0.1" $env:PORT = "8080" uvx arxiv-mcp-server --storage-path C:\absolute\path\to\papers
{ "mcpServers": { "arxiv": { "type": "http", "url": "http://127.0.0.1:8080/mcp" } } }

Cloud and load-balancer probes should GEThttp://<host>:<port>/healthz. It returns200with bodyokonce the HTTP server is listening. There is no separate/readycheck: if the process is up, it is ready. The stdio transport has no HTTP endpoints.

The server binds to127.0.0.1by default and enables MCP DNS-rebinding protection. If a reverse proxy exposes the server, keep the process on a private interface and provide authentication and network controls upstream. UseALLOWED_HOSTSandALLOWED_ORIGINSfor the host and origin values forwarded by the proxy.

Environment variable names are case-insensitive through Pydantic settings.--storage-pathis a command-line option rather than an environment setting.

Paper text and LaTeX are untrusted external content. A paper can contain text intended to manipulate an AI client into ignoring its instructions or calling unrelated tools.

- Do not treat instructions found inside a paper as trusted commands.
- Use client approval controls for shell, browser, filesystem, and messaging tools.
- Review generated summaries before taking external actions.
- Keep Streamable HTTP private unless authentication is provided upstream.

SeeSECURITY.mdfor the reporting policy and threat details.

git clone https://github.com/blazickjp/arxiv-mcp-server.git cd arxiv-mcp-server uv sync --extra test --extra dev uv run pytest uv run black --check .

Run the development checkout from an MCP client with:

{ "mcpServers": { "arxiv-dev": { "command": "uv", "args": [ "--directory", "/absolute/path/to/arxiv-mcp-server", "run", "arxiv-mcp-server" ] } } }

Contributions are welcome. ReadCONTRIBUTING.mdbefore opening a pull request, and useGitHub Issuesfor reproducible bugs or scoped feature proposals.

Search global news using natural language. Webz.io News Search API returns the most relevant articles and content, with filters for source, country, language, date, sentiment, and category.

Verified, tier-0 regulatory data for your AI: connect Claude, ChatGPT or Cursor to 850+ official sources across 50+ jurisdictions.

A flexible service for searching and analyzing academic papers on arXiv.

Read Australian Commonwealth law as it stood on any date back to 1901, and verify statute citations against the official Federal Register of Legislation. No API key.

Search scientific papers with structured experimental data extracted from full-text studies. Returns 25+ fields per paper including methods, results, sample sizes, limitations, and quality scores.

Search academic references from arXiv, DBLP, Semantic Scholar, and OpenAlex, and generate BibTeX entries.

Search and access academic paper metadata from Crossref.

Search and cite exact passages across complete classical and world-literature corpora.

Anonymous, read-only, source-backed Buddhist scripture search, passage guidance, explanation, and one-time practice planning through four production MCP tools.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.

Videos about ArXiv

Relevant YouTube tutorials, setups, and demos