Server data from the Official MCP Registry
Plain-English training-data scout: 12 experts sweep HuggingFace + GitHub, compose scored packs.
About
Plain-English training-data scout: 12 experts sweep HuggingFace + GitHub, compose scored packs.
Security Report
Valid MCP server (2 strong, 4 medium validity signals). No known CVEs in dependencies. Package registry verified. Imported from the Official MCP Registry.
5 files analyzed · 1 issue found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Required
This plugin requests these system permissions. Most are normal for its category.
How to Install
Add this to your MCP configuration file:
{
"mcpServers": {
"io-github-the-warl0ck-harvest-mcp": {
"args": [
"-y",
"@the_warl0ck/harvest-mcp"
],
"command": "npx"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
harvest-mcp
Plain-English topic in → 12 experts fan out over Hugging Face + GitHub in parallel → coverage-scored training pack JSON out in ~30 seconds.
No API keys. No Bridge. No account. Works out of the box once connected to any MCP host — no LLM inside; your agent brings the thinking.
Harvest is a training-data scout: you describe what you want to train on in plain English
("image classification", "medical question answering"), and 12 domain experts —
code, math, science, language, vision, audio, medical, law, knowledge, safety, affect, systems —
each sweep Hugging Face datasets and GitHub repos with their own curated queries. The local mixer
then composes the hits into a balanced, coverage-scored training library pack
(flare-harvest-library v1 JSON) you can hand to any training pipeline.
Install
npx github:The-Warl0ck/harvest-mcp
Or add it to Claude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"harvest": {
"command": "npx",
"args": ["github:The-Warl0ck/harvest-mcp"]
}
}
}
With optional tokens for higher rate limits:
{
"mcpServers": {
"harvest": {
"command": "npx",
"args": ["github:The-Warl0ck/harvest-mcp"],
"env": {
"HF_TOKEN": "hf_...",
"GITHUB_TOKEN": "ghp_..."
}
}
}
}
Requirements: Node.js ≥ 20.
Example session
You: find me training data for a small vision-language model
harvest.search
{ "topic": "vision language instruction tuning" }→ ~40 hits: LLaVA-NeXT data, the Cauldron, DOCCI, LLaVA repo… plus the 12-expert atlas sweep for breadth (code, math, safety…)
harvest.compose
{ "goal": "small vision-language model", "catalog": […] }→{ "title": "small vision-language model", "picks": [ { "id": "lmms-lab/LLaVA-NeXT-Data", "expert": "vision", "why": "Anchor Vision on Hugging Face" }, … ], "rationale": "Local mixer staged 22 items across 12/12 experts. No mouth required.", "engine": "local" }
harvest.pack
{ "goal": "small vision-language model", "name": "VLM Mix v1", "catalog": […] }→flare-harvest-libraryv1 JSON withcoverage: { filled: 12, totalExperts: 12, balance: 0.94 }and per-itemingest: { type: "hf-dataset" | "github-repo", ref }pointers.
The second half — the hands. Packs start as pointers; pull turns them into files:
harvest.resolve
{ "item": { "id": "openai/gsm8k", "source": "huggingface" } }→ the loot preview: 4 parquet shards + README, each with a one-line reason, nothing stupid
harvest.pull
{ "pack": {…}, "destination": "download" }→ resolves + fetches every item, updates the pack'sindexstanza, and buildsVLM Mix v1.harvest.zip— one file with everything plusMANIFEST.jsonprovenance
harvest.push
{ "bundle": "/path/to/VLM Mix v1.harvest.zip", "owner": "you", "repo": "my-data", "token": "ghp_…", "create_if_missing": true }→ one commit on a new private repo: files + manifest, message says what it is and where it came from
The same loop works for code. Second session — the build pack:
You: I want the rate-limiting code from a few good repos, packaged so my agent can build with it
harvest.search
{ "topic": "express rate limiting middleware" }→ thecodeexpert surfaces the repos with the parts you want
harvest.pack
{ "goal": "add rate-limiting to my API", "name": "Rate Limit Parts v1", "catalog": […], "purpose": "code" }→flare-harvest-codepackv1 JSON — same item envelope, no training-coverage scoring
harvest.pull
{ "pack": {…}, "destination": "download", "kinds": ["code", "docs"] }→ pulls just the source files (skips weights/datasets), builds the.harvest.zip
harvest.push
{ "bundle": "…", "owner": "you", "repo": "rate-limit-parts", "create_if_missing": true }→ your agent pulls the repo and builds from the parts
Tools
| Tool | What it does |
|---|---|
harvest.search | {topic, hf_token?, gh_token?} — topic search + full 12-expert atlas sweep, deduplicated |
harvest.expert | {expert, hf_token?, gh_token?} — run one expert's curated queries (12 ids: code, math, science, language, vision, audio, medical, law, knowledge, safety, affect, systems) |
harvest.latest | {} — trending: recently-updated HF datasets, hot/recent GitHub repos |
harvest.compose | {goal, catalog} — local mixer composes a balanced set; picks + rationale (the mixer itself needs no LLM) |
harvest.pack | {goal, catalog, name, purpose?: "training" | "code"} — compose → pack JSON. training (default) scores the mix and emits flare-harvest-library v1; code emits a flare-harvest-codepack v1: a parts bin an agent pulls and builds from |
harvest.resolve | {item, kinds?, hf_token?, gh_token?} — pull-list preview for one pack item: which files, why, and needs_token for gated items. kinds (e.g. ["code","docs"]) previews only those file kinds. No downloading |
harvest.pull | {pack, destination: "local" | "download", kinds?, dir?, confirm_large?, hf_token?, gh_token?} — resolve + fetch + index update for the whole pack; "download" also builds a .harvest.zip. kinds (e.g. ["code","docs"]) is the code-pack flow: just the source files, no weights/datasets. Oversized pulls return needs_confirm first |
harvest.push | {bundle, pack_name?, owner, repo, token, create_if_missing?, branch?, path?} — push a bundle to GitHub in one commit (private by default when creating) |
catalog items are { id, source, expert, description? } — the objects harvest.search returns drop straight in.
Rate limits
Everything hits the public Hugging Face and GitHub APIs. Unauthenticated, you get:
- Hugging Face: generous anonymous quota on
/api/datasets(occasional 429s on bursts; results are cached 12 min in-process) - GitHub: 60 requests/hour per IP for search — the 12-expert sweep is the hungriest call
For real use, set HF_TOKEN / GITHUB_TOKEN (free accounts) via env or the per-call
hf_token / gh_token args. See .env.example.
Pack format
Packs are flare-harvest-library v1 JSON — see PACK-FORMAT.md.
Pointer vs. pulled
A fresh pack is pointers, not data: each item names its source and how to
ingest it, but no files. harvest.pull turns pointers into files and records
the result in an additive index stanza on the pack JSON:
{ "index": { "openai/gsm8k": { "local": "/tmp/harvest-pulls/pack/openai__gsm8k",
"url": "https://huggingface.co/datasets/openai/gsm8k", "kind": "dataset",
"license": null, "fetched_at": "2026-09-26T…", "files": ["main/train-….parquet", …] } } }
v1 readers ignore the stanza safely; packs without it load as pointer-only.
destination: "download" additionally produces a .harvest.zip containing the
files plus MANIFEST.json (source, URL, license, SHA-256 per file — licenses
are reported as unknown when the pack doesn't state one, never guessed).
Token story
Tokens are passed per call, never stored:
harvest.search/harvest.expert/harvest.latest/harvest.resolve/harvest.pullaccepthf_token/gh_token(or theHF_TOKEN/GITHUB_TOKENenv vars) — used for higher rate limits and gated datasets, sent only as anAuthorizationheader.harvest.pushtakes a per-calltoken(GitHub PAT, orGITHUB_TOKENenv).- No token is ever written to disk, to logs, to a pack file, or to a manifest.
Fetch resume ledgers (
.harvest-fetch.json) hold only paths, byte counts, and hashes.
How it works
src/harvest/scan-core.ts— parallel HF/GitHub scanning, 12-min result cachesrc/experts.ts— the 12-expert atlas (curated queries + seed items)src/harvest/compose-local.ts— the local mixer ("No mouth required"): anchors one item per expert, then balances HF/GitHub sources, then fills for coveragesrc/coverage.ts— entropy-based coverage/balance scoring across the 12 expertssrc/harvest/compose-mouth.ts— opt-in: compose via any OpenAI-compatible/v1endpoint (LM Studio, etc.) instead of the local mixersrc/harvest/resolve.ts— the "which files" intelligence: HF Hub / GitHub tree listings classified into data, weights, config, code, docs (pure functions + thin API clients)src/harvest/fetch.ts— the hands: streaming downloads with resume, size guards, progress, SHA-256 per file; token arrives as a function arg, never storedsrc/harvest/bundle.ts— zips a fetched dir +MANIFEST.jsonprovenance into a.harvest.zipsrc/harvest/push.ts— pushes a bundle to GitHub via the Git Data API: one commit, private by default (// TODO(GitLab)seam sketched)src/harvest/index.ts— the agent side: additiveindexstanza +resolvePointer/getItemFiles/queryIndexmcp/server.ts— the MCP server (stdio, low-level SDK API, hand-written schemas)
License
Apache-2.0 — see LICENSE.
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Git
Freeby Modelcontextprotocol · Developer Tools
Read, search, and manipulate Git repositories programmatically
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Worldmonitor
Freeby Koala73 · Developer Tools
Live markets, conflicts, country risk, chokepoints, energy, and China decision signals. 89 tools.
Paperclip
Freeby Paperclipai · Developer Tools
Trending hip-hop artist momentum scores across four cultural dimensions.
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
