Back to Browse

Plant Genomics MCP Server

Developer ToolsModerate5.2MCP RegistryLocal
Free

Server data from the Official MCP Registry

Plant genomics MCP — 56 tools across 23 backends with cross-source synthesis.

About

Plant genomics MCP — 56 tools across 23 backends with cross-source synthesis.

Security Report

5.2
Moderate5.2Moderate Risk

Plant-genomics-mcp is a well-structured MCP server that provides 56 bioinformatics tools for plant genomics research across 23 public data sources. The server has strong code quality, proper auth mechanisms for HTTP transport, and appropriate permission scoping for its stated purpose. No critical vulnerabilities or malicious patterns detected. Minor findings relate to HTTP token enforcement, dependency pinning, and one request-context edge case. Supply chain analysis found 5 known vulnerabilities in dependencies (0 critical, 2 high severity). Package verification found 1 issue.

3 files analyzed · 10 issues found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

HTTP Network Access

Connects to external APIs or services over the internet.

env_vars

Check that this permission is expected for this type of plugin.

File System Write

Writes or modifies files on your machine. Check that this is expected for the tool.

cache_local

Check that this permission is expected for this type of plugin.

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-musharna-plant-genomics-mcp": {
      "args": [
        "plant-genomics-mcp"
      ],
      "command": "uvx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

🌱 plant-genomics-mcp

56 tools for plant-genomics locus lookup over the Model Context Protocol — 29 single-locus + 1 motif lookup + 1 region query + 1 assembly listing + 1 release report + 1 variant annotator + 1 gene-set enrichment + 1 BLAST search + 2 member lists (family entry, gene tree) + 13 parallel-batch + 5 cross-source synthesis variants. 23 free, public sources: Ensembl Plants, Phytozome BioMart, UniProtKB, Europe PMC, QuickGO, Planteome, PlantCyc/PMN, g:Profiler, AlphaFold DB, PDBe, InterPro, JASPAR, PANTHER, OrthoDB, AraGWAS, 1001 Genomes, NCBI BLAST, Gramene, KEGG, STRING-DB, ATTED-II, ThaleMine, and BAR (Bio-Analytic Resource for Plant Biology).

PyPI CI Docker Python License Glama DOI

📦 Install

# Zero-install — uv fetches and runs it on demand
claude mcp add plant-genomics --scope local -- uvx plant-genomics-mcp
# pipx — installs the CLI onto your PATH
pipx install plant-genomics-mcp
claude mcp add plant-genomics --scope local -- plant-genomics-mcp

# GHCR Docker image
docker pull ghcr.io/musharna/plant-genomics-mcp:latest
claude mcp add plant-genomics --scope local -- \
  docker run --rm -i ghcr.io/musharna/plant-genomics-mcp:latest

# From source
git clone https://github.com/musharna/plant-genomics-mcp.git
cd plant-genomics-mcp
python -m venv .venv && .venv/bin/pip install -e .
claude mcp add plant-genomics --scope local -- "$(pwd)/.venv/bin/plant-genomics-mcp"

💬 Try it

Once connected, ask Claude a plain-language question — you don't have to name any tool or remember the chain:

"Tell me everything about the Arabidopsis gene AT1G01010 — its function, GO terms, KEGG pathways, protein-interaction partners, and recent papers."

The server supplies the tools (here: Ensembl Plants, UniProt, QuickGO, KEGG, STRING-DB and Europe PMC lookups); which ones get called, in what order and in how many turns is up to the client. In the recording at the top of this page (a narrower prompt), Claude Code picked the calls itself and returned one combined answer. Swap in any locus and pass organism= for cross-species — e.g. rice Os01g0100100 (oryza_sativa) — and each tool maps the organism to that backend's own identifier.

A worked 64-call run over 114 genes in three organisms, with the 51 gaps it logged (39 since closed), is in examples/arf_family/PAGE.md.

🛠️ Tools

56 tools across 23 backends — Ensembl Plants, Phytozome BioMart, UniProtKB, Europe PMC, QuickGO, Planteome, PlantCyc/PMN, g:Profiler, AlphaFold DB, PDBe, InterPro, JASPAR, PANTHER, OrthoDB, AraGWAS, 1001 Genomes, NCBI BLAST, Gramene, KEGG, STRING-DB, ATTED-II, ThaleMine, BAR. 29 single-locus + 1 motif lookup + 1 region query + 1 assembly listing + 1 release report + 1 variant annotator + 1 gene-set enrichment + 1 BLAST search + 2 member lists + 13 parallel-batch + 5 cross-source synthesis. Most take a TAIR-style locus (e.g. AT1G01010) plus optional organism= (slug / scientific name / common name / NCBI taxid — 12-plant curated coverage matrix at the pgmcp://organisms/coverage MCP resource). All publish JSON outputSchema, EDAM ontology tags, and behaviour annotations — every tool is readOnlyHint + openWorldHint, so hosts can surface them without a destructive-action confirmation prompt.

#CategoryToolWhat it does
1Gene metadata (live)ensembl_plants_lookup_locusFetches gene record from Ensembl Plants REST (any plant species).
2Cross-references (live)get_gene_xrefsFetches cross-DB references (UniProt, NCBI Gene, TAIR, GO, …) from Ensembl.
3Gene metadata (live)phytozome_lookup_locusFetches gene record from Phytozome BioMart (any Phytozome proteome).
4Protein (live)resolve_locus_to_uniprotResolves a locus to its UniProtKB record (Swiss-Prot preferred, TrEMBL OK).
5Literature (live)locus_literatureSearches Europe PMC for papers mentioning the locus (free, no API key).
6GO annotations (live)locus_go_annotationsFetches QuickGO GO annotations (locus → UniProt → QuickGO).
7Sequence search (live)blast_sequenceNCBI BLAST URLAPI — async Put/Get polling with progress notifications.
8Homology (live)gramene_homologsFetches Gramene v69 homology entries (ortholog / paralog) with gene_tree_id.
9Pathways (live)kegg_pathwaysFetches KEGG pathway memberships. 7 organisms: Arabidopsis (ath:, native AGI), + rice (osa:), maize (zma:), soybean (gmx:), barley (hvg:), poplar (pop:), brachypodium (bdi:) bridged via Ensembl → Entrez ID.
10Interactions (live)string_interactionsFetches STRING-DB first-neighbor interaction partners with per-channel score.
11Coexpression (live)atted_coexpressionFetches ATTED-II top-N coexpression neighbors with a score: z-score for Arabidopsis, logit score (LSmr) for the other releases.
12Curator summary (live)bar_gene_summaryFetches BAR ThaleMine + GAIA-aliases curator summary for an Arabidopsis locus.
13Expression (live)bar_efp_expressionFetches BAR eFP-Browser expression profile (mean ± SD per tissue) for a locus.
14Interactions (live)bar_aiv_interactionsFetches BAR AIV interaction partners (Arabidopsis + rice) with confidence + papers.
15Curator summary (live)tair_locus_infoSilent upgrade — alias of bar_gene_summary. MCP tool name preserved for clients.
16Metabolism (live)plantcyc_locus_infoWalks gene → enzyme → reactions → PlantCyc/PMN pathways (free BioCyc web-services API). The metabolic-pathway view KEGG/GO lack; found=false for non-enzymatic genes. 11 species have a PGDB.
17Sequence (live)get_sequenceFetches a locus's sequence (genomic / cds / cdna / protein) from Ensembl /sequence/id — the fetch half of lookup → fetch → BLAST; feed sequence to blast_sequence.
18Region query (live)ensembl_region_queryLists gene/transcript/cds/exon features overlapping a genomic interval (chr:start-end) via Ensembl /overlap/region — "what's in this QTL interval" without a per-locus lookup.
19Enrichment (live)go_enrichmentGO + KEGG over-representation for a gene list via g:Profiler g:GOSt — "what is my DE / co-expression set enriched for?" Reports unmapped loci; optional custom background. All 12 organisms.
20Plant ontology (live)locus_plant_ontologyPlant Ontology (anatomy / dev-stage) + Trait Ontology annotations for a locus via Planteome (Solr) — the plant-specific ontologies GO doesn't cover. by_ontology rollup; taxon-filtered. Strong for 6 species.
21Structure (live)alphafold_structureAlphaFold DB predicted 3D model for a locus (locus → UniProt → model): global mean pLDDT, per-band confidence, modelled span, and mmCIF / PDB / PAE URLs. found=false when no model is deposited. All 12 organisms.
22Structure (live)experimental_structuresPDBe experimentally-solved (X-ray / cryo-EM / NMR) structures for a locus (locus → UniProt): best-first PDB id, chain, method, resolution, coverage, residue span. found=false when none deposited (common for plants). All 12 organisms.
23Domains (live)interpro_domainsInterPro domain / family architecture (locus → UniProt): each entry's accession, name, type, source_database (Pfam included), integrated InterPro id, and residue spans, plus a count_by_type rollup. All 12 organisms.
24TF motifs (live)tf_binding_motifsJASPAR curated TF DNA-binding profiles for a locus (locus → UniProt → symbol search, then UniProt-confirmed): matrix id, TF class/family, assay type (SELEX / ChIP-seq / PBM / DAP-seq), IUPAC consensus, PubMed refs, logo URL. Fuzzy name hits for other genes are quarantined in name_only_matches. Arabidopsis-heavy coverage.
25TF motifs (live)jaspar_motifOne JASPAR profile by matrix id (e.g. MA0570.1, or MA0570 for the newest version) including the raw position-frequency matrix — the drill-down companion to tf_binding_motifs.
26Interactions (live)experimental_interactionsThaleMine CURATED EXPERIMENTAL interaction partners (BioGRID / IntAct / PSI-MI) for an Arabidopsis locus — per partner: detection method (two hybrid, pull down, ...), PSI-MI relationship type, physical vs genetic, source DB, PubMed IDs, and an evidence count. The experimental counterpart to string_interactions (predicted / text-mined). Arabidopsis only.
27Function (live)locus_gene_rifsThaleMine curated GeneRIF statements — one-sentence, manually curated descriptions of what the gene does, each tied to a PubMed ID (HY5 has 114). Citable functional context that GO terms and raw abstracts don't provide. Arabidopsis only.
28Variation (live)locus_variantsNatural (EVA/dbSNP) variants overlapping a locus's genomic span via Ensembl /overlap/region — id, source, consequence class, alleles, clinical significance. variant_count + truncated. All 12 organisms.
29Variation (live)vep_annotateEnsembl VEP consequence prediction for a variant (region + allele, not locus) — most-severe consequence + per-transcript SO terms, IMPACT, SIFT. All 12 organisms.
30Orthology (live)panther_familyPANTHER protein family + subfamily (id + name), GO terms by aspect, protein class, and pathways. found=false when unmapped. All 12 organisms.
31Orthology (live)orthodb_orthologsOrthoDB ortholog group (name, evolutionary rate) + cross-species member genes at the Viridiplantae level. organism_count + truncated. All 12 organisms.
32Diversity (live)aragwas_associationsAraGWAS genome-wide association hits per locus — score, MAF, SNP effect, phenotype/study. Arabidopsis-only.
33Diversity (live)arabidopsis_natural_variation1001 Genomes natural-variation SNP effects across 1135 accessions — chr, position, effect, impact, amino-acid change, transcript + gene span. Arabidopsis-only.
34Batch (live)batch_* (twelve dedicated variants)Parallel per-locus fanout for tools 1–6, 8–12, 14. Up to 50 loci per call.
35Batch (live)batch_locus_callRuns any tool whose only required argument is locus (33 tools, including gene_report and the synthesis tools) over up to 50 loci; the shared args are checked against that tool's schema once, before any call.
36Synthesis (live)*_synth / consensus_homologs (four)Compose 2–5 backends in parallel, return a SynthesisEnvelope with per-step status.
37Synthesis (live)gene_reportOne-shot "tell me about this gene" dossier — annotation + xrefs + protein + domains + GO + KEGG + STRING + literature composed into a rendered Markdown result.markdown (+ structured result.sections).
38Families (live)entry_membersEvery protein in one organism carrying an InterPro / Pfam / PANTHER entry, with the locus each maps to (entry → genes; the reverse of tools 23 and 30). UniProt total + cursor paging; reviewed-only by default.
39Families (live)gene_tree_membersEvery gene in an Ensembl Compara gene tree — the gene_tree_id that gramene_homologs returns — with locus, protein id, species and organism; target_organism keeps one organism's members. total + limit.
40Homology (live)ensembl_plants_paralogsParalogues Ensembl Compara records for a locus — within_species_paralog and the other_paralog ("ancient paralogues") that gramene_homologs drops — closest first, with perc_id, taxonomy level and protein id. Not a family list: test membership with interpro_domains. total + limit.
41Region query (live)ensembl_plants_assemblyAn organism's Ensembl assembly — name, GCA accession, karyotype — and every top-level seq-region with its length: the region names ensembl_region_query takes, chromosomes first, so a region walk is planned before its first call. total + limit.
42Provenance (live)upstream_releaseThe release a backend's own endpoint calls current (Ensembl, STRING, QuickGO, JASPAR, KEGG), for the backends whose answers state none; PDBe, AraGWAS and Europe PMC publish none and say why. A separate request, so read it before and after a run: equal values mean no release changed.

⚡ Quickstart

After install, the simplest call returns the Ensembl Plants record for NAC001 — the canonical worked example used throughout examples/:

// arguments
{ "locus": "AT1G01010" }

// result (truncated)
{
  "id": "AT1G01010",
  "organism": "arabidopsis_thaliana",
  "display_name": "NAC001",
  "biotype": "protein_coding",
  "seq_region_name": "1",
  "start": 3631,
  "end": 5899,
  "strand": 1,
  "assembly_name": "TAIR10",
  "description": "NAC domain containing protein 1 ..."
}

Cross-species — pass organism=:

{ "locus": "Os01g0100100", "organism": "oryza_sativa" }

A recorded Claude Code session (2026-05-24) with a narrower prompt — the Ensembl record, UniProtKB entry and top three Europe PMC papers for AT1G01010 — answered it in one turn:

Full per-tool walkthroughs (with real upstream-API transcripts) live in examples/:

WalkthroughCoverage
gene_report_AT1G01010.mdOne-shot Markdown gene dossier — 7 backends composed, with graceful KEGG degradation.
analyze_locus_AT1G01010.mdEnsembl → xrefs → UniProt → Europe PMC → QuickGO chain (5 tools).
find_homologs_AT1G01010_NAC_domain.mdBLAST + per-hit UniProt enrichment.
biological_context_AT1G01010.mdGramene + KEGG + UniProt + STRING + ATTED-II (5 tools).
v0.8_synthesis_walkthrough.mdThe four v0.8 synthesis tools (*_synth + consensus_homologs) on one locus. Captured 2026-05-22 at v0.8.
cross_organism_walkthrough.mdv0.9 multi-organism resolver against rice + maize — per-backend routing. Captured 2026-05-24 against PyPI v1.0.4.

📚 Resources & prompts

Clients discover them via resources/list and prompts/list.

Resources (resources/read):

URIWhat
pgmcp://cache/statsPer-backend TTLCache rollup — {hits, misses, size} for each live backend.
pgmcp://organisms/phytozomeSlug → Phytozome organism_id map.
pgmcp://backends/statusStatic per-backend roster — name, base_url, citation DOI, subscription_gated, kind. Nothing is probed: kind is always "live" and says nothing about whether the backend is up right now.
pgmcp://organisms/coverageMarkdown table of all 12 supported plants × 9 ID slots (ncbi_taxid / ensembl / phytozome / string / europe_pmc / kegg / atted / gprofiler / plantcyc).

Prompts (prompts/get):

NameRequiredOptionalChains
analyze_locuslocusorganism (default arabidopsis_thaliana)Ensembl → xrefs → UniProt → Europe PMC → QuickGO.
find_homologssequenceprogram (default blastp)blast_sequence → per-hit resolve_locus_to_uniprot for UniProt-shaped accessions.
biological_contextlocustop_n (default 10)Gramene → KEGG → UniProt → STRING → ATTED-II.

🔌 Transports

TransportHow to launch
stdio (default)plant-genomics-mcp (after install) or via Docker above
streamable-HTTPplant-genomics-mcp-http — POST JSON-RPC at http://host:port/mcp

The HTTP transport is stateless and emits JSON responses by default — the right shape for registry indexers and remote hosting.

Self-hosting

There is no public hosted endpoint. To use the HTTP transport, run it yourself: it is the same binary, gated by your own bearer token (PLANT_GENOMICS_MCP_HTTP_TOKEN), with NCBI BLAST requests sent under your own contact email.

⚙️ Configuration

Stdio needs no configuration. The two env vars that matter:

VariableWhenEffect
PLANT_GENOMICS_MCP_HTTP_TOKENHTTP transport onlyBearer token for /mcp; must be ≥32 chars or the HTTP server aborts at startup. Generate openssl rand -hex 32.
PLANT_GENOMICS_MCP_NCBI_EMAILIf you use BLASTNCBI etiquette contact. Unset → placeholder + per-call warning; NCBI may throttle.
VariableDefaultEffect
PLANT_GENOMICS_MCP_HTTP_HOST127.0.0.1HTTP bind address.
PLANT_GENOMICS_MCP_HTTP_PORT8765HTTP TCP port.
PLANT_GENOMICS_MCP_HTTP_MAX_BODY2097152 (2 MiB)Reject POSTs with Content-Length larger than this.
PLANT_GENOMICS_MCP_HTTP_STATELESS10 keeps per-client session state (SSE-style).
PLANT_GENOMICS_MCP_HTTP_JSON10 switches the response shape to streaming SSE events.
PLANT_GENOMICS_MCP_BLAST_CONCURRENCY2Max in-flight BLAST searches per process (NCBI per-IP rate limit).
PLANT_GENOMICS_MCP_CACHE_TTL600Per-backend TTL+LRU cache entry lifetime, in seconds. 200-only.
PLANT_GENOMICS_MCP_CACHE_SIZE256Max entries per backend before LRU eviction.
PLANT_GENOMICS_MCP_CACHE_DISABLEDunsetAny non-empty value makes every cache a no-op.

The cache is process-local — restart the server to drop all entries. Long-running calls (retry storms, multi-second Phytozome BioMart POSTs) emit MCP notifications/progress over the active session; clients opt in via progressToken in the request _meta.

⚠️ Error model

All live tools raise PlantGenomicsError subclasses; the MCP SDK stringifies them into the wire content with a [ClassName] prefix so clients can route on failure kind without parsing the message:

Wire prefixWhen
[NotFoundError]404 / empty BioMart row / invalid locus identifier
[RateLimitError]429 retry budget exhausted — back off and retry
[UpstreamUnavailableError]5xx past retry budget — service outage, try a peer backend
[PlantGenomicsError]Other (BioMart Query ERROR: body, unexpected column count, etc.)

Batch tools return {tool, count, results, errors} where results[locus] is the same shape as the single-locus tool and errors[locus] is the same [ClassName] message string. Ensembl's batch uses the native POST /lookup/id endpoint (one HTTP round-trip); everything else fans out via asyncio.gather.

🧪 Development

.venv/bin/pip install -e '.[dev]'                         # or: uv sync --extra dev
.venv/bin/pytest -q                                       # unit tests
PLANT_GENOMICS_MCP_LIVE=1 .venv/bin/pytest -q             # adds live network probes
PLANT_GENOMICS_MCP_STDIO_SMOKE=1 .venv/bin/pytest -q      # adds stdio smoke
.venv/bin/ruff check .

With uv, pass --extra dev — a bare uv sync omits (and removes) the test dependencies. See CONTRIBUTING.md.

CI runs the unit suite + the stdio smoke on every push/PR (matrix: Python 3.11, 3.12, 3.13, 3.14 — the full requires-python range), with a dead proxy set so an ungated network call fails. A separate live-smoke job runs two live test files (tests/test_verify_genes.py, tests/test_arf_mcp_client.py, InterPro and PANTHER) on Python 3.12 on the same pushes and PRs; it reports a result but is not a required check, because its result depends on third-party services. The rest of the live suite (PLANT_GENOMICS_MCP_LIVE=1) is not run in CI.

Drift detection. scripts/benchmark_annotations.py runs a curated corpus of 27 loci (spanning all 12 organisms) through 12 functions — the organism resolver plus 11 lookups across 9 of the 23 backends (ATTED-II, BAR, Ensembl Plants, Europe PMC, Gramene, KEGG, Phytozome, STRING-DB, UniProt) — and compares the results to a frozen snapshot of earlier results (scripts/benchmark_annotations.expected.json), emitting PASS / DRIFT / FAIL plus two cross-source consistency invariants. A change is detected relative to that snapshot, not checked against an independent truth. BLAST and the synthesis pipelines are registered in the script, but no corpus locus exercises them. A scheduled GitHub Actions workflow (.github/workflows/benchmark.yml) runs it weekly and pages when the same loci fail on a re-run. Operator guide: docs/benchmarking.md.

.venv/bin/python scripts/benchmark_annotations.py        # full live sweep

See CHANGELOG.md for release notes, including the v0.8 → v0.9 species=/organism_id= → organism= migration and the v1.0.1 HTTP-token enforcement change.

MCP registry

Listed in the official MCP registry under the namespace below (ownership-verification token for mcp-publisher):

mcp-name: io.github.musharna/plant-genomics-mcp

License

MIT — see LICENSE. Underlying services (Ensembl Plants, Phytozome, TAIR, PlantCyc, BAR) have their own terms of use; consult each before bulk querying.

Reviews

No reviews yet

Be the first to review this server!