Server data from the Official MCP Registry
LLM evals as MCP tools: score outputs for faithfulness, relevancy, and hallucination.
About
LLM evals as MCP tools: score outputs for faithfulness, relevancy, and hallucination.
Security Report
retriEVAL is a well-structured MCP evaluation server with appropriate authentication, reasonable permissions, and no critical vulnerabilities. The codebase demonstrates good security practices for its use case (LLM evaluation with judge backends). Minor concerns exist around input validation in the sandbox route and verbose error messages, but these are low-severity and do not significantly impact security posture. Supply chain analysis found 5 known vulnerabilities in dependencies (0 critical, 5 high severity).
4 files analyzed · 11 issues found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Required
This plugin requests these system permissions. Most are normal for its category.
What You'll Need
Set these up before or after installing:
Environment variable: ANTHROPIC_API_KEY
Environment variable: RETRIEVAL_JUDGE_BACKEND
Environment variable: RETRIEVAL_JUDGE_MODEL
Environment variable: RETRIEVAL_BUDGET_USD
How to Install
Add this to your MCP configuration file:
{
"mcpServers": {
"io-github-hcarrillo001-retrieval-mcp": {
"env": {
"ANTHROPIC_API_KEY": "your-anthropic-api-key-here",
"RETRIEVAL_BUDGET_USD": "your-retrieval-budget-usd-here",
"RETRIEVAL_JUDGE_MODEL": "your-retrieval-judge-model-here",
"RETRIEVAL_JUDGE_BACKEND": "your-retrieval-judge-backend-here"
},
"args": [
"-y",
"github:hcarrillo001/retrieval-mcp"
],
"command": "npx"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
retriEVAL
LLM evaluation as an MCP server. Score your AI's outputs for faithfulness, relevancy, and hallucination from inside any MCP client — no pipeline, no test harness. Every result comes back with a link to a dashboard that keeps the history.
Try it live (no signup) · Watch the 2-minute demo · Dashboard
Why
Five customer-support answers, scored on two metrics:
| metric | score | passing |
|---|---|---|
| answer_relevancy | 0.98 | 5/5 |
| faithfulness | 0.70 | 3/5 |
Every answer was on-topic and well-written. Two of them contradicted the policy they were supposedly grounded in — one promised free return shipping the policy doesn't offer, another invented a free overnight replacement. Reviewing by eye, you'd sign off on all five.
That gap is the point. Relevancy asks did it answer the question. Faithfulness asks is it actually in the source. You need both, and the second one catches the expensive failures.
Connect
Self-hosted. Clone it, point it at a judge, run it — your data never leaves your machine and there's no service to sign up for.
git clone https://github.com/hcarrillo001/retrieval-mcp
cd retrieval-mcp
pip install -r requirements.txt
export ANTHROPIC_API_KEY=sk-ant-... # or a local judge, below
python server.py # stdio, for Claude Desktop / Cursor
Then just ask:
Score these cases with faithfulness: [{"input": "...", "actual_output": "...", "retrieval_context": ["..."]}]
Pass your cases inline and nothing is stored — one call, no setup step. See Run locally (stdio) for client config, and Deploy as HTTP if you want your own always-on instance with a dashboard.
Want to try it before installing anything? There's a live sandbox at retrieval-mcp.com — no signup, runs on a free judge, nothing saved.
Keeping everything local: set RETRIEVAL_JUDGE_BACKEND=ollama and the judge
runs on your machine too, so no data leaves your network at any point. Useful if
you're evaluating anything you can't send to a third party.
What you get
- 9 built-in metrics plus custom metrics you author in plain English
- Swappable judges — Anthropic, Groq, Gemini, OpenRouter, or a local Ollama model, so nothing has to leave your network
- Golden sets from files, URLs, inline JSON, JSONL, CSV, or TSV
- Run history in Supabase with shareable permalinks and run comparison
- A spend cap, because a judge-based tool can otherwise run up a bill
Honest limitations
- Judge agreement hasn't been validated against human labels yet, so treat scores as a signal rather than ground truth.
- Golden sets currently hold their own outputs, so re-running one against new model outputs means loading a second set. Splitting them is the next change.
- Golden sets and authored metrics live in the server process and are lost on restart. Runs persist; those don't.
Metrics (DeepEval-aligned)
faithfulness · answer_relevancy · contextual_precision ·
contextual_recall · contextual_relevancy · hallucination · bias ·
toxicity · summarization — plus authored G-Eval metrics you define in
plain language. All are normalized so higher = better (bias/toxicity report
the clean fraction), and each reasons before scoring.
Versatile golden sets
load_golden_set accepts a file path (including uploaded files), an
http(s) URL, an inline JSON array, or JSONL text, in
JSON / JSONL / CSV / TSV. Field names are auto-normalized (question→input,
answer→actual_output, ground_truth→expected_output, contexts→context,
passages→retrieval_context, …), so most public benchmarks load as-is.
Short replies by default
run_eval scores every case but returns only the 3 lowest-scoring by
default (tune with limit), with total_cases/shown and a pointer to
show_run_cases(run_id, offset, limit, metric) to page through the rest.
Spend cap (so it can't run up a bill)
Judge spend is metered from real token usage and persisted. Set a hard cap:
export RETRIEVAL_BUDGET_USD=20 # 0/unset = unlimited
Once cumulative spend hits the cap, further Anthropic calls stop and tools
return a clear budget_exceeded message. Check/clear with get_budget /
reset_budget. (Prices are approximate — override RETRIEVAL_PRICE_IN/OUT
$/1M tokens to match current pricing for your model.)
Hybrid setup (local + always-on dashboard)
Run history uses a pluggable store, chosen by env:
- FileStore (default) — JSONL in
~/.retrieval. Zero setup, local only. - SupabaseStore — when
SUPABASE_URL+SUPABASE_SERVICE_KEYare set. Run history lives in Postgres, shared by the local CLI, the deployed MCP, and the website dashboard.
Recommended hybrid flow:
- Run
supabase_schema.sqlin Supabase (creates therunstable). - Set
SUPABASE_URL+SUPABASE_SERVICE_KEYon the MCP (local and/or Railway) so every run is written centrally. Each run records itsgenerator_modelandjudge_modelfor cross-model comparison. - Deploy
web/to Vercel (set the same Supabase env vars) and map it toretrieval-mcp.com. The dashboard reads history via/api/runs(service key stays server-side) and renders trend-by-model, a model leaderboard, and run history. It shows sample data until Supabase is wired.
Local stays your free sandbox (Ollama judge, file history); the website is the always-on window into the shared history.
Run locally (stdio) — Claude Desktop
pip install -r requirements.txt
export ANTHROPIC_API_KEY=sk-ant-...
{
"mcpServers": {
"retrieval": {
"command": "python",
"args": ["/ABSOLUTE/PATH/server.py"],
"env": { "ANTHROPIC_API_KEY": "sk-ant-...", "RETRIEVAL_BUDGET_USD": "10" }
}
}
}
Then: "Load examples/rag_golden.jsonl as 'space', run faithfulness, label it v1."
Deploy as HTTP (reach it from anywhere)
export RETRIEVAL_TOKEN=$(openssl rand -hex 24) # required for a public endpoint
export ANTHROPIC_API_KEY=sk-ant-...
export RETRIEVAL_BUDGET_USD=20
python app.py # serves $PORT (default 8000); MCP at /mcp, health at /healthz
Deploy to Railway (or any host): the included Dockerfile / Procfile
work as-is. Set ANTHROPIC_API_KEY, RETRIEVAL_TOKEN, RETRIEVAL_BUDGET_USD
in the host env. Clients connect to https://<host>/mcp with header
Authorization: Bearer <token> — add it as a custom connector in
claude.ai / Claude Desktop, or point Agent Builder / CI at it. State (golden
sets, run history, spend) lives server-side, so it persists across machines.
See it on a real RAG pipeline (demo)
demo/rag_demo.py builds a tiny end-to-end RAG over a small labeled dataset
(demo/labeled.json + demo/corpus.json): it retrieves with BM25, computes
recall@k against the gold passages (deterministic — the retriever's score),
generates an answer, then scores faithfulness (the generator's score). One answer
is deliberately hallucinated so you watch the two failure modes separate.
python demo/rag_demo.py # offline, no key needed
python demo/rag_demo.py --real # real generation + RetriEval judge (needs ANTHROPIC_API_KEY)
It also writes demo/generated_goldenset.jsonl — load that into the MCP
(load_golden_set → run_eval) for the judge-scored version. This is the bridge:
your pipeline emits predictions, the dataset supplies the labels, and RetriEval
scores retriever and generator independently.
Connecting to a RAG pipeline
- Offline (default): export your pipeline's retrieved context + answer into a golden set and score it — RetriEval never touches your pipeline.
- Live: add a
query_rag(question)tool that calls your RAG endpoint or your vector store (Chroma / Supabase pgvector), captures context + answer, and scores in one shot.
Judge backend
export RETRIEVAL_JUDGE_BACKEND=anthropic # default
export RETRIEVAL_JUDGE_MODEL=claude-sonnet-4-6
# or local, free:
export RETRIEVAL_JUDGE_BACKEND=ollama
export RETRIEVAL_JUDGE_MODEL=deepseek-r1:70b
Jev (experimental)
Jev is a decision model: it returns probabilities,
not text. With RETRIEVAL_JUDGE_BACKEND=jev, faithfulness becomes a hybrid:
an LLM splits the answer into standalone claims, Jev checks every claim against
the context in parallel, and the score is supported claims / total claims
(the same definition as the LLM rubric). The reason lists the unsupported claims
with their probabilities, and details.claims has every claim's p_supported.
hallucination uses the same claims with one Jev Choice per claim (supports /
contradicts / says nothing): score = 1 - contradicted / total, and claims the
context is silent on are reported as "not in context" rather than penalised.
answer_relevancy is a Jev Score on a 5-level rubric and needs no LLM call.
Each claim also carries the exact passage of the answer it came from (quote),
which the sandbox highlights. Every other metric still runs on the LLM set by
RETRIEVAL_DECOMPOSER_BACKEND.
export RETRIEVAL_JUDGE_BACKEND=jev
export TYPESAFE_API_KEY=... # console.typesafe.ai/keys
export RETRIEVAL_DECOMPOSER_BACKEND=anthropic # or ollama / openai
# optional: JEV_MODEL (jev-latest), JEV_MAX_WORKERS (8), JEV_PRICE_IN (0.042)
Or leave the server default alone and pick Jev per call: run_eval and
evaluate_case take judge="jev" (or judge="llm" to force the LLM judge).
Once TYPESAFE_API_KEY is set on the server, the sandbox picker also lists
"Jev · TypeSafe", with claims split by the Groq preset.
Jev spend counts toward RETRIEVAL_BUDGET_USD. Jev is weak on numbers, dates
and adversarial text, and its agreement with human labels has not been measured
here yet, so treat Jev scores as experimental.
Reports: why a run passed or failed
Every run_eval result carries aggregate[metric].verdict: a headline ("Fails:
faithfulness averaged 0.33, below the 0.70 threshold") and the drivers behind
it (the failing cases, or the unsupported claims). It is saved with the run.
- In Claude,
summary_mdincludes the verdict, and clients that support MCP Apps render a compact card inline: the headline score, the other metrics, and a checklist of what failed and why. - The full report is
report_url(/report?run=<id>on your dashboard host): score, verdict, every claim, and the answer with its problem passages highlighted. It is the same view the sandbox shows after a run, drawn by one shared renderer (web/report.js,web/report.css).
Tools
| Tool | Purpose |
|---|---|
list_metrics | built-in + authored metrics |
load_golden_set(name, source, fmt) | name a set for reuse (self-host only — shared and lost on restart) |
list_golden_sets | what's loaded |
author_metric(name, criteria, examples) | plain language → a scorer |
run_eval(metrics, cases, golden_set, threshold, outputs, label, limit) | score a set; pass cases inline (JSON/JSONL/CSV/TSV/path/URL) — nothing stored |
show_run_cases(run_id, offset, limit, metric) | page the rest |
evaluate_case(...) | one-off score |
ground_against_url(url, output, question) | check an output's consistency with a web page (no labels — consistency, not correctness) |
list_runs(golden_set, last_n) | saved runs |
plot_metric_trend / plot_run / compare_runs | inline charts |
get_budget / reset_budget | spend cap status / reset |
License
Apache License 2.0. Built by Hanns Carrillo.
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Git
Freeby Modelcontextprotocol · Developer Tools
Read, search, and manipulate Git repositories programmatically
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Worldmonitor
Freeby Koala73 · Developer Tools
Live markets, conflicts, country risk, chokepoints, energy, and China decision signals. 90 tools.
Paperclip
Freeby Paperclipai · Developer Tools
Trending hip-hop artist momentum scores across four cultural dimensions.
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
