Back to Browse

Bioharbor MCP Server

Developer ToolsModerate7.3MCP RegistryLocal
Free

Server data from the Official MCP Registry

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

About

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

Security Report

7.3
Moderate7.3Low Risk

BioHarbor is a well-structured bioinformatics MCP server with solid security fundamentals. Authentication and token storage are properly handled through environment variables. Code quality is good with appropriate input validation and error handling. Permissions are appropriately scoped to the server's stated purpose (bioinformatics compute, GPU management, file I/O). Minor code quality issues and some unvalidated subprocess arguments do not significantly impact the overall security posture. Package verification found 1 issue.

6 files analyzed Β· 6 issues found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

HTTP Network Access

Connects to external APIs or services over the internet.

File System Read

Reads files on your machine. Normal for tools that analyze or process local data.

File System Write

Writes or modifies files on your machine. Check that this is expected for the tool.

env_vars

Check that this permission is expected for this type of plugin.

Shell Command Execution

Runs commands on your machine. Be cautious β€” only use if you trust this plugin.

gpu_control

Check that this permission is expected for this type of plugin.

system_info

Check that this permission is expected for this type of plugin.

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-danielluo2-bioharbor": {
      "args": [
        "bioharbor"
      ],
      "command": "uvx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

BioHarbor

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

CI PyPI License

πŸ§ͺ Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome β€” see the roadmap.

BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools β€” not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.

BioHarbor demo

Why another bio MCP server?

Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them β€” BioHarbor complements them by running the compute:

Database MCP serversBioHarbor
Runs real analyses (search, fold, cluster)βŒβœ…
Validates inputs before burning GPU timeβŒβœ…
GPU-aware queue, polite on shared GPUsβŒβœ…
Long jobs return a job_id instead of timing outβŒβœ…
Compact summaries + files on disk (saves tokens)βŒβœ…
Provenance for every run, export to a pipelineβŒβœ… (export: planned)

Quick start

pip install bioharbor
bioharbor doctor             # checks Python, GPUs, workspace, tools
bioharbor setup-db swissprot # reference database for search_homologs (needs MMseqs2)
bioharbor install            # shows how to connect Claude, Cursor or Codex

For structure prediction on a GPU: pip install "bioharbor[esmfold]" β€” see docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).

Connect your agent

BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):

ClientSet up
Claude Codeclaude mcp add bioharbor -- bioharbor serve
Claude Desktopbioharbor install claude-desktop --write, then restart the app
Cursorbioharbor install cursor --write, or Add to Cursor
Codex (CLI, IDE extension, app)codex mcp add bioharbor -- bioharbor serve, or bioharbor install codex --write
Biomni (Stanford's biomedical agent)agent.add_mcp(...); see docs/use-with-biomni.md
Anything elserun bioharbor serve (stdio) or bioharbor serve --http (Streamable HTTP)

Long-running tools return a job_id within ~20 s instead of blocking, so they stay within every client's tool-call timeout.

Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.

Then ask your agent something like:

Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?

Use it without an agent

Every tool is also a CLI command, with identical behaviour:

bioharbor tools list
bioharbor run find_orfs sequence=@contig.fa min_aa=100 --brief   # human-readable
bioharbor run seq_stats sequence=MKTAYIAKQRQISFVKSHFSRQ
bioharbor jobs

Shared GPU server

GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:

# on the GPU server
bioharbor serve --http --host 127.0.0.1 --port 8765
# on your laptop, then point Cursor / Codex / Claude Code at http://127.0.0.1:8765/mcp
ssh -N -L 8765:127.0.0.1:8765 you@gpu-server

⚠️ HTTP mode has no authentication yet (on the roadmap), so keep it on 127.0.0.1 and use the tunnel. Details: docs/connect-clients.md.

Tools

ToolWhat it doesRuns
seq_statsValidate sequences; type, length, GC%, molecular weightinline
translate_sequenceDNA/RNA β†’ protein, one or all six framesinline
find_orfsLongest ORFs on both strands, with coordinatesinline
search_homologsMMseqs2 search (protein, or translated DNA) vs local DBsjob
predict_structureESMFold structure, pLDDT bands, low-confidence regions, pTMjob (GPU)
scrna_pipelinescanpy QC β†’ clustering β†’ markersplanned

Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status, read_file.

How it works

Agent ──MCP──▢ validate input ─▢ inline? ──yes──▢ run ─┐
                                   β”‚ no                  β”œβ”€β–Ά provenance + summary ─▢ Agent
                                   β–Ό                     β”‚
                     job queue (SQLite) ─▢ GPU placement β”˜
                     (waits politely for a GPU with free memory)
  • Every call is a job recorded in SQLite with params, versions, timings and GPU used, plus a provenance.json next to its outputs.
  • GPU placement reads live free memory and utilisation (NVML or nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected.
  • Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
  • Results are agent-shaped: summary, message, files, suggestions. Errors carry a hint and a retryable flag.

Details: docs/design.md.

Writing a tool

from pydantic import BaseModel, Field
from bioharbor.registry import Resources, RunContext, tool
from bioharbor.results import ToolResult


class FoldParams(BaseModel):
    sequence: str = Field(..., description="Protein sequence")


@tool(
    version="1",
    slow=True,
    resources=Resources(gpu=True, gpu_mem_gb=lambda p: 4 + len(p.sequence) / 100),
)
def predict_structure(params: FoldParams, ctx: RunContext) -> ToolResult:
    """Predict a protein structure with ESMFold."""
    ...
    return ToolResult(summary={"mean_plddt": 87.1}, files=["model.pdb"])

Plugins can ship tools in their own package via the bioharbor.tools entry-point group. See CONTRIBUTING.md.

License

Apache-2.0

Reviews

No reviews yet

Be the first to review this server!