Back to Browse

Open Access API Harvester MCP Server

Developer ToolsUse Caution4.2MCP RegistryLocal
Free

Server data from the Official MCP Registry

Every OpenAlex search stores its request, response and checksums so sources can be checked later.

About

Every OpenAlex search stores its request, response and checksums so sources can be checked later.

Security Report

4.2
Use Caution4.2High Risk

Notanda is a scholarly document harvesting tool with a local MCP server component. The codebase demonstrates reasonable security practices including proper credential handling, input escaping in the web UI, and use of safe libraries (defusedxml). However, there are some code quality concerns around error handling, overly broad exception catches, and the MCP server's authentication model relies entirely on environment variables with no token validation. Supply chain analysis found 5 known vulnerabilities in dependencies (0 critical, 5 high severity). Package verification found 1 issue (1 critical, 0 high severity).

4 files analyzed · 12 issues found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

File System Read

Reads files on your machine. Normal for tools that analyze or process local data.

File System Write

Writes or modifies files on your machine. Check that this is expected for the tool.

HTTP Network Access

Connects to external APIs or services over the internet.

env_vars

Check that this permission is expected for this type of plugin.

system_info

Check that this permission is expected for this type of plugin.

Unverified package source

We couldn't verify that the installable package matches the reviewed source code. Proceed with caution.

What You'll Need

Set these up before or after installing:

OpenAlex API key used for literature searches.Required

Environment variable: HARVESTER_OPENALEX_API_KEY

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-rudiregenwurm-notanda": {
      "env": {
        "HARVESTER_OPENALEX_API_KEY": "your-harvester-openalex-api-key-here"
      },
      "args": [
        "notanda"
      ],
      "command": "uvx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

Notanda — Open-Access API Harvester

Notanda is the public product name and PyPI distribution name. The repository, Python import package and CLI keep their established technical names Open-Access-API-Harvester, harvester and harvester.

This repository publishes the reviewed 1.3.2 product-core source. It is not a hosted service and the recommended accompanied beta path remains beta@notanda.io.

A public Installer Preview is available under GitHub Releases solely as a technical pre-release: https://github.com/RudiRegenwurm/Open-Access-API-Harvester/releases/tag/v1.3.0-installer-preview.1 Those installers are unsigned on Windows/Linux and ad-hoc signed on macOS, may trigger SmartScreen/Gatekeeper warnings, and are not the recommended installation path for non-technical users.

A headless, resumable, idempotent CLI pipeline that discovers Open-Access scholarly works through OpenAlex Topics, cross-checks them against Europe PMC, falls back to Unpaywall for OA locations, validates every downloaded artifact, and stores the result as a deterministic flat corpus ready for a downstream pre-ingestion pipeline. Schema 4 additionally retains provider observations, merge decisions and acquisitions in an append-only Evidence Ledger while the established read model remains compatible.

DISCOVERY → NORMALIZATION → DEDUPLICATION → ACQUISITION → VALIDATION → INGESTION → PROVENANCE/REPORTING

It is not a web application, document-management system, search engine, OCR system or LLM application. It is a document harvesting pipeline with persistent state.


Quick start

With Python 3.10 or newer, install the PyPI distribution in a fresh virtual environment:

python -m venv .venv
# Linux/macOS:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install notanda
harvester --version

For an isolated application install with pipx:

pipx install notanda

For development from a repository checkout:

python3 -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"

With the local control center (no command line needed after this)

harvester serve

Opens a local web UI at http://localhost:8765/ for harvesting, browsing the corpus, inspecting runs and provenance, verifying integrity, retrying failures and editing settings. Credentials can be entered in its Settings page on first run.

New Harvest offers two search modes. Conventional Search uses the query you type, exactly as typed. Assisted Search turns a plain-language research question into one compact retrieval query you can edit, previews up to ten real discovery results without downloading anything, and only then hands the approved query to the same harvest pipeline. Assisted Search needs a query-advisor credential; Conventional Search does not and is unaffected when the advisor is missing or unavailable.

Experimental local MCP server

The repository also contains an optional local MCP component under mcp/. It exposes two deliberately narrow tools — search_literature and get_evidence — and persists the request, provider observation, exact returned result and SHA-256 manifest under one stable evidence ID. It performs no full-text download or corpus write and opens no network server.

The separate notanda-mcp Beta package uses the reviewed notanda==1.3.1 PyPI release as its core. With Python 3.11 or 3.12, install it using python -m pip install "notanda-mcp==0.1.0b2". The MCP Registry entry describes the package, local stdio transport, required OpenAlex API key and writable evidence directory. The package is not included in the notanda 1.3.2 distribution. The technical live search and restart/receipt checks passed on 2026-10-02; Thomas's external unaccompanied test is still pending. Client configuration, evidence format and verification steps are in mcp/README.md.

From the command line

export HARVESTER_OPENALEX_API_KEY='…'          # free: openalex.org/settings/api
export HARVESTER_CONTACT_EMAIL='ops@example.org'  # required by Unpaywall
export HARVESTER_ADVISOR_API_KEY='…'           # optional: Assisted Search in the UI only

harvester discover --topic-id T10159 --limit 25    # dry run, downloads nothing
harvester harvest  --topic-id T10159 --limit 25    # real harvest
harvester verify --deep                            # check the corpus against state

Both drive the same engine and the same state database: a run started in the browser can be resumed from the terminal.

Output

The downstream contract is a flat directory of deterministic sibling files:

<storage_root>/
├── doi_10_1371_journal_pone_0123456_5d41402abc4b.pdf     primary full text
├── doi_10_1371_journal_pone_0123456_5d41402abc4b.xml     when legitimately available
└── doi_10_1371_journal_pone_0123456_5d41402abc4b.json    mandatory sidecar

The sidecar keeps bibliographic metadata, artifact metadata, provenance and harvest state in separate blocks. Missing information is null; nothing is ever invented.

{
  "document_id": "doi_10_1371_journal_pone_0123456_5d41402abc4b",
  "doi": "10.1371/journal.pone.0123456",
  "title": "…",
  "abstract": "…",
  "oa_status": "gold",
  "oa_status_source": "openalex",
  "domain_tags": ["Social Sciences", "Psychology", "…"],
  "artifacts": { "pdf": { "sha256": "…", "size_bytes": 123456, "…": "…" } },
  "provenance": {
    "discovered_via": ["openalex"],
    "cross_checked_via": ["europe_pmc"],
    "acquired_via": "unpaywall",
    "resolved_url": "…",
    "http_status": 200
  },
  "harvest": { "run_id": "…", "status": "COMPLETED", "attempts": 1 }
}

What it guarantees

GuaranteeHow
Nothing false succeedsHTTP 200 is never enough: magic bytes, trailer, parser openability, size bounds and SHA-256 must all pass before an artifact exists under its final name.
Atomic artifactsEvery download goes to a unique .part file and is published with a single atomic rename after validation.
ResumableDiscovery cursors and document state are checkpointed continuously; harvester resume continues without repeating completed work.
IdempotentRe-running the same harvest downloads nothing and produces a byte-identical corpus.
DeduplicatedOne logical document per canonical DOI, however many providers report it.
No silent lossEvery failure is a structured row in the state database and appears in the run report.
Historical evidenceRe-observations, field-level merge decisions and acquisitions append immutable rows; current documents, source_records and artifacts remain compatible projections.
Budget-safeProvider daily-budget exhaustion checkpoints and suspends cleanly (exit 4) instead of hammering the API.
Secret-safeAPI keys and contact addresses are redacted from logs, provenance, reports and state.
LawfulOnly OA locations that a provider explicitly flags as open are fetched. Access controls are recorded, never circumvented.

Commands

CommandPurpose
serveopen the local web control center
harvestdiscover + acquire
discoverdiscovery and normalization only (dry run)
resumecontinue an interrupted or suspended run
statusrun and corpus status (--json)
inspectone document's canonical record (--json)
verifycheck the corpus against state (--deep, --json)
retry-failedre-queue failed documents
evidence-exportwrite a deterministic, independently readable ledger JSON bundle (--output)
evidence-restorevalidate and restore a bundle into an empty ledger (--input)

The evidence bundle contains Ledger records and file references, not the PDF/XML artifacts or the complete corpus directory. A portable corpus export with files is still in development.

Exit codes: 0 success · 1 failure · 2 configuration error · 3 completed with document failures · 4 suspended (provider budget) · 130 interrupted.

Tests

pytest            # offline suite — no network, no credentials required
pytest -m live    # controlled real-provider smoke tests

The offline suite runs against deterministic in-process mock providers covering normal responses, pagination, throttling, budget exhaustion, timeouts, malformed payloads, missing full text, alternative locations and duplicates.

Documentation

FileContents
LICENSEMIT license for the product code and documentation, subject to the stated exceptions
THIRD_PARTY_NOTICES.mdbundled font licenses and separately licensed dependencies
TRADEMARKS.mdtreatment of the Notanda name and visual identity assets
CODE_SIGNING_POLICY.mdWindows code signing policy for the SignPath Foundation application
INSTALL-MACOS.mdmacOS Installer Preview and Gatekeeper exception procedure
mcp/README.mdexperimental local MCP server, evidence contract and verification guide

Code signing policy

Windows installers are currently unsigned.

Dependencies

httpx (streaming HTTP with a pluggable transport), pypdf (PDF structural validation), defusedxml (safe XML parsing). Everything else is the standard library: configuration, SQLite persistence, CLI, logging, hashing, concurrency, retry — and the web UI, which uses http.server plus a front end with no build step, so harvester serve is the only command needed to run it.

License

Unless a file or notice says otherwise, the product source code and documentation in this repository are licensed under the MIT License, copyright 2026 Rudolf Kiechle.

The bundled IBM Plex and Source Serif 4 font files are not MIT-licensed; they remain under the SIL Open Font License 1.1 included beside the files. External Python dependencies are not vendored and retain their own licenses. The Notanda name, emblem, wordmark, lockup, favicons and avatar are excluded from the MIT grant. Details and exact file scopes are in THIRD_PARTY_NOTICES.md and TRADEMARKS.md.

Reviews

No reviews yet

Be the first to review this server!