Unofficial Ensembl MCP Server

SSE

by Augmented-Nature

173 downloads Not rated yet MIT license

About

A comprehensive Model Context Protocol (MCP) server that provides access to the Ensembl REST API for genomic data, comparative genomics, and biological annotations.

Details

Transport
SSE
License
MIT license

Explore

- Regulatory Elements: Access enhancers, promoters, and TFBS data
- Motif Features: Get transcription factor binding motifs
- Cell Type Context: Filter regulatory features by cell type

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Unofficial Ensembl MCP Server
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository


npm install

1. Build the server (if not already done):

bash
npm run build
``

2. Add to Claude Desktop configuration:

- Open Claude Desktop
- Go to Settings → MCP Servers
- Add a new server with:
- Name:
ensembl
- Command:
node
- Args:
/path/to/ensembl-server/build/index.js

3. Restart Claude Desktop to load the server

Once connected, you can use natural language to access genomic data:

- "Look up the BRCA2 gene and get its sequence"
- "Find orthologs of TP53 in mouse"
- "Get variants in the region chr17:43044295-43125364"
- "Search for insulin-related genes"
- "Get the assembly information for human genome"
- "Translate this DNA sequence to protein: ATGAAACGC..."

- ENSEMBL_BASE_URL: Override default API base URL
-
REQUEST_TIMEOUT: Set custom timeout (default: 30000ms)

- Default species: homo_sapiens`
- Automatic species validation
- Support for all Ensembl divisions

ensembl_list_species

List species supported by Ensembl with display name, common name, assembly, taxon ID, and division. Required discovery step — species names like homo_sapiens are opaque to non-biologists and are the input format every other Ensembl tool expects. Filter by division to select one; use nameContains to find a species by partial name match. With no division, returns the endpoint default division — the vertebrates (~356 species on the default GRCh38 endpoint); pass a division to list that division.

ensembl_lookup_gene

Resolve a gene by symbol + species (or by stable ID) to its Ensembl ID, genomic location (chr:start-end:strand), biotype, description, and transcript list. Entry point for most workflows — the stable ID and coordinates returned here are inputs to other tools. Accepts both symbol lookup (BRCA2 + homo_sapiens) and direct ID lookup (ENSG00000139618). Supports batch lookup of up to 20 IDs or symbols in one call via the ids or symbols field. Provide exactly one of symbol, id, ids, or symbols. For …

ensembl_get_sequence

Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region. Returns the sequence with its stable ID, molecule type, and character count — large sequences are returned in full but the length is stated so callers can budget context. The type parameter selects which sequence is fetched: genomic (default, includes introns), cdna (spliced transcript), cds (coding sequence only), protein. For region mode, set id to a region — either species:chr:start-end (e.g. …

ensembl_query_region

Find genomic features overlapping a chromosomal region: genes, transcripts, variants, regulatory elements, or exons. Returns each feature with its stable ID, type, location, biotype, and name. Useful for "what's in this locus?" and for seeding follow-up lookups. Region format is chr:start-end (e.g. 13:32315086-32400268 for the BRCA2 locus). Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted.…

ensembl_predict_variant

Predict the functional consequences of a sequence variant using the Ensembl Variant Effect Predictor (VEP). Accepts three input formats: HGVS notation (transcript-relative, e.g. ENST00000380152.8:c.2T>A, or genomic, e.g. 13:g.32316462T>A); region+allele (chr:start:end:strand/allele, e.g. 1:65568:65568:1/T); and a dbSNP rsID (e.g. rs334). Returns the most severe consequence term, affected transcripts and genes, impact level (HIGH/MODERATE/LOW/MODIFIER), and any colocated known variants with cl…

ensembl_get_homology

Find orthologs and/or paralogs of a gene across species. Returns each homolog's stable ID, species, homology type (ortholog_one2one, ortholog_one2many, paralog_many2many, etc.), perc_id (percent identity), perc_pos (percent positives), and taxonomy level. Essential for cross-species research — for example, "what is the mouse equivalent of human TP53?" or "how conserved is BRCA2 across mammals?". Provide either symbol + species or a stable gene ID. Target species can be filtered to a single sp…

ensembl_get_xrefs

Retrieve cross-database references for a gene or feature — HGNC, UniProt, EntrezGene, OMIM, RefSeq, Reactome, and others. Returns each xref with its database name, primary ID, display ID, and description. The dbname filter narrows to specific databases; omit to return all xrefs. IDs returned here chain to protein (pubchem via UniProt), literature (pubmed via PubMed IDs), disease (OMIM via MIM_GENE), and pathway (Reactome) resources. Requires an Ensembl stable ID — use ensembl_lookup_gene to g…

This server integrates well with other bioinformatics MCP servers:

- UniProt Server: Protein data integration
- AlphaFold Server: 3D structure predictions
- STRING Server: Protein interaction networks
- PDB Server: Structural biology data

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "unofficial ensembl mcp server": {
            "ensembl": {
                "command": "node",
                "args": [
                    "/path/to/ensembl-server/build/index.js"
                ]
            }
        }
    }
}

McpServers

{
    "ensembl": {
        "command": "node",
        "args": [
            "/path/to/ensembl-server/build/index.js"
        ]
    }
}

A comprehensive Model Context Protocol (MCP) server that provides access to the Ensembl REST API for genomic data, comparative genomics, and biological annotations.

Developed by Augmented Nature

Overview

This server enables seamless access to Ensembl's vast genomic database through a standardized MCP interface. It supports gene lookups, sequence retrieval, variant analysis, comparative genomics, regulatory features, and much more across multiple species.

Features

Gene & Transcript Information

- Gene Lookup: Get detailed gene information by Ensembl ID or gene symbol
- Transcript Analysis: Retrieve all transcripts for a gene with structural details
- Gene Search: Search genes by name, description, or identifier with filtering options

Sequence Data

- Genomic Sequences: Extract DNA sequences for any genomic region or feature
- CDS Sequences: Get coding sequences for specific transcripts
- Sequence Translation: Translate DNA sequences to protein sequences
- Repeat Masking: Support for hard and soft repeat masking

Comparative Genomics

- Homolog Detection: Find orthologous and paralogous genes across species
- Phylogenetic Trees: Generate gene family trees in multiple formats
- Cross-Species Analysis: Compare genes and genomes across different organisms

Variant Data

- Variant Retrieval: Get genetic variants in genomic regions
- Consequence Prediction: Predict variant effects on genes and transcripts
- Population Genetics: Access allele frequencies and population data

Regulatory Features

- Regulatory Elements: Access enhancers, promoters, and TFBS data
- Motif Features: Get transcription factor binding motifs
- Cell Type Context: Filter regulatory features by cell type

Cross-References & Annotations

- External Database Links: Get cross-references to PDB, EMBL, RefSeq, etc.
- Coordinate Mapping: Convert coordinates between genome assemblies
- Ontology Terms: Access GO terms and functional annotations

Species & Assembly Information

- Species Lists: Browse available species and assemblies
- Assembly Statistics: Get genome assembly information and statistics
- Karyotype Data: Access chromosome information and banding patterns

Batch Processing

- Batch Gene Lookup: Process multiple genes simultaneously
- Batch Sequence Fetch: Retrieve sequences for multiple regions efficiently

Installation

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.