PubChem MCP Server
About
Provides comprehensive access to PubChem's chemical information database via the PubChem PUG REST API.
Details
- Transport
- SSE
Explore
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
PubChem MCP ServerCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
pubchem_search_compounds
Search PubChem for chemical compounds by identifier (name, SMILES, or InChIKey, batched up to 25), molecular formula in Hill notation, substructure or superstructure containment, or 2D Tanimoto similarity. Optionally hydrate results with properties to avoid a follow-up pubchem_get_compound_details call.
pubchem_get_compound_details
Get detailed compound information by CID. Returns physicochemical properties (molecular weight, SMILES, InChIKey, XLogP, TPSA, etc.), optionally with a textual description (pharmacology, mechanism, therapeutic use), all known synonyms, drug-likeness assessment (Lipinski/Veber rules), and/or pharmacological classification (FDA classes, MeSH classes, ATC codes). Efficiently batches up to 100 CIDs.
pubchem_get_compound_image
Fetch a 2D structure diagram (PNG image) for a compound by CID.
pubchem_get_compound_3d_structure
Get a compound's default 3D conformer — atomic coordinates and bonds — for one CID. format="json" (default) returns parsed atoms and bonds the model can reason over directly; format="sdf" returns the raw V2000 SDF text for passthrough to docking, rendering, or conformer tools. Optionally lists alternate conformer IDs. Not every compound has computed 3D coordinates (large molecules, mixtures, and some salts do not).
pubchem_get_compound_xrefs
Get external database cross-references for a compound: PubMed citations, patent IDs, gene/protein associations, registry numbers, and taxonomy IDs. Results are capped per type with total counts reported.
pubchem_get_compound_safety
Get GHS (Globally Harmonized System) hazard classification and safety data for one or more compounds by CID. Returns signal word, pictograms, hazard statements (H-codes), and precautionary statements (P-codes) per compound. Data sourced from PubChem depositors — source attribution included.
pubchem_get_bioactivity
Get a compound's bioactivity profile: which assays tested it, activity outcomes (Active/Inactive/Inconclusive), target identifiers (NCBI Gene ID, UniProt/GenBank accession), and quantitative values (IC50, EC50, Ki, etc.). Filter by outcome and/or a specific molecular target (NCBI Gene ID or protein accession) to focus the profile — e.g. "is this compound active against target T?".
pubchem_get_compound_interactions
Get a compound's interaction data: drug-drug interactions (DrugBank), drug-food interactions, and chemical-target interactions (binding/activity from BindingDB, ChEMBL, and others). Each entry carries its originating source. Richest for approved drugs; many compounds have no deposited interaction records.
pubchem_search_assays
Find PubChem bioassays associated with a biological target. Search by gene symbol (e.g. "EGFR"), protein name, NCBI Gene ID, or UniProt accession. Returns assay IDs (AIDs) which can be explored further with pubchem_get_summary.
pubchem_get_summary
Get descriptive summaries for PubChem entities by ID. Supports assays (AID), genes (Gene ID), proteins (UniProt accession), and taxonomy (Tax ID). Up to 10 per call.
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"pubchem mcp server": {
"server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/pubchem-mcp-server",
"run",
"start:stdio"
],
"env": {
"MCP_LOG_LEVEL": ""
}
}
}
}
}
McpServers
{
"server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/pubchem-mcp-server",
"run",
"start:stdio"
],
"env": {
"MCP_LOG_LEVEL": ""
}
}
}
Transport
"stdio"
Package
"@cyanheads/pubchem-mcp-server"
Registry
"npm"
Provides comprehensive access to PubChem's chemical information database via the PubChem PUG REST API.
Search the PubChem chemical database for compounds, properties, safety data, bioactivity, cross-references, and entity summaries via MCP. STDIO or Streamable HTTP.
Public Hosted Server:https://pubchem.caseyjhand.com/mcp
Ten tools for querying PubChem's chemical information database:
Search PubChem for chemical compounds across five search modes.
- Identifier lookup— resolve compound names, SMILES, or InChIKeys to CIDs (batch up to 25)
- Formula search— find compounds by molecular formula in Hill notation
- Substructure/superstructure— find compounds containing or contained within a query structure
- 2D similarity— find structurally similar compounds by Tanimoto similarity (configurable threshold)
- Caps at 200 CIDs per page;offsetpages further, to a ceiling of 10,000. Identifier lookups page over the set already resolved; formula and structure searches widen their bounded upstream request to reach a page, so deep pages cost more upstream
- Optionally hydrate results with properties to avoid a follow-up details call
Get detailed compound information by CID.
- Batches up to 100 CIDs in a single request
- 27 available properties: molecular weight, SMILES, InChIKey, XLogP, TPSA, complexity, stereo counts, and more
- Optionally includes textual descriptions (pharmacology, mechanism, therapeutic use) from PUG View — fetched for the first 10 CIDs of a batch, with the skipped CIDs named in the response
- Optionally includes known synonyms (trade names, systematic names, registry numbers)
- Synonyms and descriptions are paged:synonymOffsetanddescriptionOffsetwindow every compound in the batch at the same position, reaching the entries past a page
- Optionally computes drug-likeness assessment (Lipinski Rule of Five + Veber rules) from fetched properties
- Optionally fetches pharmacological classification (FDA classes, mechanisms of action, MeSH classes, ATC codes)
Get a compound's bioactivity profile from PubChem BioAssay.
- Returns assay outcomes (Active/Inactive/Inconclusive), target info (protein accessions, NCBI Gene IDs), and quantitative values (IC50, EC50, Ki)
- Filter by outcome and/or a specific molecular target (NCBI Gene ID or protein accession)
- Caps at 100 results per page;offsetreaches the rest (well-studied compounds may have thousands)
Get descriptive summaries for four PubChem entity types.
- Assays (AID), genes (Gene ID), proteins (UniProt accession), taxonomy (Tax ID)
- Up to 10 entities per call
- Type-specific field extraction for clean, structured output
Get a compound's interaction data by CID.
- Drug-drug interactions (DrugBank), drug-food interactions, and chemical-target binding/activity (BindingDB, ChEMBL, and others)
- Select which interaction kinds to fetch and cap entries per kind
- Paged per kind: each reports its source-record total and its ownnextOffset, andoffsetreaches the records past a page
- Each entry carries its originating source — coverage is richest for approved drugs
Get a compound's default 3D conformer by CID.
- format="json"returns parsed atoms (element + x/y/z) and bonds for direct reasoning;format="sdf"returns raw V2000 SDF for passthrough to docking or rendering
- maxAtoms/maxBondsbound the atom/bond preview andincludeRawSdfopts into a large raw SDF past the safe line cap;atomCount/bondCountalways report the totals and any capping is disclosed
- Optionally lists alternate conformer IDs
- Returns a typed not-found when PubChem has no computed 3D coordinates (large molecules, mixtures, some salts)
Compound and assay records are also exposed as URI-templated MCP resources, backed by the same client methods as the tools:
- Declarative tool definitions — single file per tool, framework handles registration and validation
- Unified error handling across all tools
- Pluggable auth (none,jwt,oauth)
- Swappable storage backends:in-memory,filesystem,Supabase,Cloudflare KV/R2/D1
- Structured logging with optional OpenTelemetry tracing
- Runs locally (stdio/HTTP) or containerized via Docker
- Rate-limited client for PUG REST and PUG View APIs (5 req/s with automatic queuing)
- Retry with exponential backoff on 5xx errors and network failures
- All tools are read-only and idempotent — no API keys required
A public instance is available athttps://pubchem.caseyjhand.com/mcp— no installation required. Point any MCP client at it via Streamable HTTP:
{ "mcpServers": { "pubchem-mcp-server": { "type": "streamable-http", "url": "https://pubchem.caseyjhand.com/mcp" } } }
Add to your MCP client config (e.g.,claude_desktop_config.json):
{ "mcpServers": { "pubchem-mcp-server": { "type": "stdio", "command": "bunx", "args": ["@cyanheads/pubchem-mcp-server@latest"], "env": { "MCP_TRANSPORT_TYPE": "stdio" } } } }
git clone https://github.com/cyanheads/pubchem-mcp-server.git
No API keys are required — PubChem's API is freely accessible.
bun run rebuild bun run start:stdio # or start:http
bun run devcheck # Lints, formats, type-checks bun run test # Runs test suite
docker build -t pubchem-mcp-server . docker run -p 3010:3010 pubchem-mcp-server
SeeCLAUDE.mdfor development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — notry/catchin tool logic
- Usectx.logfor domain-specific logging
- Register new tools in theindex.tsbarrel file
Issues and pull requests are welcome. Run checks before submitting:
Access and interact with Allen Institute for Neural Dynamics (AIND) metadata directly within your IDE.
A high-performance JavaScript server for the Alliance of Genome Resources (AGR) MCP.
Access the AlphaFold Protein Structure Database for protein structure prediction and analysis.
Interface with Biomart, a biological data query tool, using the pybiomart Python package.
Agent-first rewrite of genomeoncology's BioMCP in TypeScript to provide next-gen biomedical data access for agents.
Perform complex queries on the DANDI Archive, a platform for neurophysiology data.
Interact with DROMA drug-omics association analysis databases using natural language.
A bridge to the Drug Gene Interaction Database (DGIdb) API, enabling AI clients to query drug-gene interaction data.
Query the Materials Project database using the mp_api client. Requires an MP_API_KEY environment variable.
Access PubMed articles through the Entrez API.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



