Server data from the Official MCP Registry
Compare text LLMs across OpenRouter, Bedrock, Vertex AI and Foundry with automated grading.
About
Compare text LLMs across OpenRouter, Bedrock, Vertex AI and Foundry with automated grading.
Security Report
EvalForge Lite is a legitimate MCP server for comparing LLMs across multiple providers. The codebase demonstrates reasonable security practices with proper credential handling (credentials not stored, env vars used, no hardcoded secrets), but has moderate concerns: credentials are sent in POST request bodies as plaintext JSON (though over HTTPS in normal deployment), the web app stores sensitive data in browser memory without explicit protection, and there is minimal input validation on custom model IDs and test prompts. Permissions appropriately match the server's purpose of calling external LLM APIs. Supply chain analysis found 14 known vulnerabilities in dependencies (0 critical, 6 high severity). Package verification found 1 issue.
4 files analyzed · 22 issues found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Required
This plugin requests these system permissions. Most are normal for its category.
What You'll Need
Set these up before or after installing:
Environment variable: OPENROUTER_API_KEY
How to Install
Add this to your MCP configuration file:
{
"mcpServers": {
"io-github-thejaredchapman-evalforge-lite": {
"env": {
"OPENROUTER_API_KEY": "your-openrouter-api-key-here"
},
"args": [
"evalforge-lite"
],
"command": "uvx"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
EvalForge Lite
Compare text LLMs side by side. Write a few test prompts, pick up to four models across OpenRouter, Amazon Bedrock, Google Vertex AI and Microsoft Foundry, and EvalForge Lite sends the same prompts to all of them, scores the answers automatically, and shows a leaderboard with letter grades, response time, speed and estimated cost. Use it as a web app or as an MCP server for Claude and other assistants. You bring your own credentials, or the person hosting it keeps them on the server for you.
Documentation site: https://thejaredchapman.github.io/evalforge-lite/
Why use it
- Test on your own prompts. Choose models from evidence about your tasks, not a generic benchmark.
- Up to 4 models per run, across 4 backends.
X(OpenRouter) andX@bedrockare separate targets, so you can check one model on two platforms in a single run. - Automatic grading. A judge model scores each answer against your rubric, and rule checks (
contains,regex,json_valid,max_length, available through the API and MCP) add a pass or fail. You get a 0-100 score and a letter grade. - A second opinion on every response. Each answer is also evaluated on six criteria (answered, quality, instruction following, completeness, helpfulness, safety) with strengths and weaknesses written out.
- Speed and cost beside quality. Latency, tokens per second and estimated cost for every model, and a "What matters most?" selector that moves the "Best for ..." badge without a new run.
- Policy gate. Upload a company policy and prompts that violate it are blocked before any model is called. If the check itself fails, the prompt is blocked.
- Reports. Download a PDF or a CSV for any of your last five runs.
- No accounts, no database. Credentials are used for one request and not stored. Nothing is written to disk.
Quick start
1. Run the web app on your computer
Requires Python 3.10 or newer (3.12 recommended; download from https://www.python.org/downloads/) and git.
git clone https://github.com/thejaredchapman/evalforge-lite.git
cd evalforge-lite
python3.12 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
python app.py
Open http://localhost:8000, paste a key for at least one backend (an OpenRouter API key is the quickest: https://openrouter.ai/workspaces/default/keys), add a test case, pick two to four models, and click Run comparison. Runs are limited to 3 per 8 hours per browser session. Full walkthrough: Getting started.
2. Use it from Claude (MCP server)
With uv installed:
uvx evalforge-lite
Add it to Claude Code in one line:
claude mcp add evalforge-lite -- uvx evalforge-lite
Or install the Claude Code plugin, which bundles the same server:
claude plugin marketplace add thejaredchapman/evalforge-lite
claude plugin install evalforge-lite@evalforge
Then ask your assistant to compare models. It gets 9 tools: list_models,
suggest_models, list_availability, set_policy, evaluate_prompt,
run_comparison, list_runs, get_report, get_report_csv.
Details, Claude Desktop config and credential shapes: MCP server.
3. Host it for other people
Deploy with the included render.yaml (gunicorn, one worker) or any host that
can run gunicorn --workers 1 --threads 4 --bind 0.0.0.0:$PORT app:app. By
default every visitor supplies their own key. Optionally keep provider keys on
the server with environment variables and a shared daily cap (50 per 24 hours by
default). Keep it at one worker: all state is in memory per process.
Full guide: Hosting and server-side keys.
Good to know
- The app does not read
.envby itself. To use values from it, runset -a; source .env; set +abeforepython app.py. - Each run is limited to 4 models, and each browser session gets 3 runs per 8 hours.
- All state lives in memory and is cleared when the server restarts. See Privacy and limits.
- Upgrading from an older version and reading the CSV or API fields? See the notes in Troubleshooting and FAQ.
Backends and credentials
| Backend | What you provide |
|---|---|
| OpenRouter | One API key |
| Amazon Bedrock | A region, plus a Bedrock API key or AWS access keys (optional session token) |
| Google Vertex AI | A project id and region, plus an access token or service-account JSON |
| Microsoft Foundry | A resource name and region, plus an API key or Entra ID access token |
A separate judge backend setting chooses where the judge and policy gate run. Bedrock, Vertex and Foundry costs are estimates from catalog prices, not your cloud bill. See Backends and credentials.
Documentation
| Page | What is in it |
|---|---|
| Overview | What it is, who it is for, the three ways to use it |
| Getting started | Install, run, and your first comparison |
| Web app guide | Every part of the screen, in order |
| Comparing models | Reading metrics, grades, evaluation, cost and their limits |
| Backends and credentials | Keys, regions, X@backend targets |
| MCP server | Install paths, all 9 tools, example prompts |
| Hosting and server-side keys | Deploying for others, operator-held keys, daily cap |
| Troubleshooting and FAQ | Common messages, fixes, and notes on CSV/API field changes |
| Privacy and limits | What data goes where, what is stored, every limit |
Test
pytest tests/ -v
Every model and HTTP call is mocked, so the suite needs no API key and makes no network calls.
Contributing
Contributions are welcome: bug reports, model-catalog updates, new checks, docs, and new backends. See CONTRIBUTING.md for setup, tests and the pull request process. When the app shows an error, the popup's Report an issue on GitHub button opens a pre-filled bug report.
License
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Git
Freeby Modelcontextprotocol · Developer Tools
Read, search, and manipulate Git repositories programmatically
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Worldmonitor
Freeby Koala73 · Developer Tools
Live markets, conflicts, country risk, chokepoints, energy, and China decision signals. 86 tools.
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
MCP Marketplace
Freeby mcp-marketplace · Developer Tools
Search and install MCP servers from inside your AI client.
