Server data from the Official MCP Registry
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
About
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
Remote endpoints: streamable-http: https://www.nonobench.com/mcp
Security Report
Valid MCP server (2 strong, 2 medium validity signals). No known CVEs in dependencies. Imported from the Official MCP Registry.
11 tools verified · Open access · No issues found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Required
This plugin requests these system permissions. Most are normal for its category.
How to Connect
Remote Plugin
No local installation needed. Your AI client connects to the remote endpoint directly.
Add this to your MCP configuration to connect:
{
"mcpServers": {
"io-github-mauricekleine-nonobench": {
"url": "https://www.nonobench.com/mcp"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
Nonobench
A benchmark suite for evaluating LLM reasoning capabilities on Nonogram (Picross) puzzle solving across different grid sizes. Results are published at nonobench.com.
Built by Maurice Kleine.
What is a Nonogram?
Nonograms (also known as Picross, Griddlers, or Paint by Numbers) are logic puzzles where you fill in cells on a grid based on numeric clues for each row and column. The clues indicate consecutive groups of filled cells, separated by at least one empty cell. Solving these puzzles requires logical deduction and constraint satisfaction - making them an excellent test of LLM reasoning abilities.
Project Structure
nonobench/
├── bench/ # Benchmark runner, results database and exporter
└── visualizer/ # TanStack Start dashboard on Cloudflare Workers (nonobench.com), also home of the puzzle set
Prerequisites
- Bun runtime (v1.4+)
- Node.js v24+ (for the visualizer)
- OpenRouter API key
Quick Start
1. Clone the Repository
git clone https://github.com/mauricekleine/nonobench.git
cd nonobench
2. Set Up Environment Variables
Create a .env file in the bench/ directory (see bench/.env.example):
OPENROUTER_API_KEY=your_openrouter_api_key_here
This is the only variable needed. The visualizer builds without any.
3. Running Benchmarks
cd bench
bun install
bun run bench # prints the plan and exits, no API calls
bun run bench --model <name> # run one model (repeat --model for more)
bun run bench --all-missing # run every configured model with missing work
bun run bench --model <name> --sizes 20x20 # opt in to the extended tier
Runs are incremental and append-only: a model/puzzle pair that already has a successful result is never run again, and the database refuses to overwrite it. Failed runs are retried on the next invocation. Results are stored in bench/results.db (SQLite).
Useful flags:
--max-cost <usd>stops launching new puzzles once this session's spend reaches the amount. Requests already in flight still finish, so a session can overshoot by up to--parallelrequests per selected model. Use the OpenRouter key's own limit as the hard ceiling.--parallel <n>sets concurrent requests per model (default 10); lower it for rate-limited providers.--limit <n>runs only the first n puzzles of each size, for pilots against a scratch database (NONOBENCH_DB=/tmp/copy.db).
Output modes
New runs ask for the answer as strict structured output (a JSON schema, only routed to endpoints that enforce it), so models cannot wrap the grid in prose. For a few models the schema-enforcing endpoints measurably hurt answers; those run in text mode instead (outputMode: "text" in bench/constants.ts, chosen by a 5x5 A/B with the benchmark prompt). Every run records its mode, and grading is identical for both: the answer must satisfy every clue. The runner also stops a model that solves none of the 5x5 puzzles with structured output, and one whose early runs mostly report zero reasoning tokens, since both point at the harness rather than the model.
After benchmarking, export results for the visualizer:
bun run export
This writes visualizer/app/results.json (aggregates) and visualizer/public/results-raw.json (every run, including prompts and outputs).
Other scripts:
bun test- parser, grader, database-policy and puzzle checksbun run typecheck- TypeScript checkbun run regrade- read-only comparison of stored grades against the current grader
Running Against a Local Model
The runner can benchmark a model that you serve yourself. Any server with an OpenAI-compatible /v1 endpoint works: vLLM, SGLang, llama.cpp, LM Studio, and Ollama.
This setup uses two machines. The model server is the workstation that serves the model. The bench machine is the laptop that runs the benchmark.
1. Serve the model on the workstation
Bind the server to all interfaces. The default bind address, 127.0.0.1, accepts connections only from the model server itself.
vLLM:
vllm serve Qwen/Qwen3-32B --host 0.0.0.0 --port 8000
Ollama:
ollama pull qwen3:32b
OLLAMA_HOST=0.0.0.0:11434 ollama serve
llama.cpp:
llama-server -m ./qwen3-32b-Q4_K_M.gguf --host 0.0.0.0 --port 8080
SGLang:
python -m sglang.launch_server --model-path Qwen/Qwen3-32B --host 0.0.0.0 --port 30000
LM Studio: open the Developer tab, then select Start Server. Turn on the Network Server switch, so the server listens on all interfaces.
CAUTION: Do not expose the port to the internet. Most local servers have no authentication, so anyone who reaches the port can run the model.
2. Open the port on the model server
Restrict the rule to your own subnet. Allow TCP traffic from your local network to the port:
sudo ufw allow from 192.168.1.0/24 to any port 8000 proto tcp
On macOS, allow incoming connections for the server app in System Settings.
3. Verify the connection from the laptop
Find the LAN address of the model server. On Linux, run ip -4 addr show. On macOS, run ipconfig getifaddr en0.
From the laptop, request the model list:
curl http://192.168.1.20:8000/v1/models
Make sure that the answer lists your model. Copy the exact id value. You need that value for NONOBENCH_LOCAL_MODEL.
If the request fails, check the bind address from step 1. Then check the firewall rule, the address, and the port.
If the model server is not reachable on your network, create an SSH tunnel to it instead. Then use the address http://127.0.0.1:8000/v1:
ssh -N -L 8000:127.0.0.1:8000 user@192.168.1.20
4. Run the benchmark
cd bench
bun install
NONOBENCH_LOCAL_BASE_URL=http://192.168.1.20:8000/v1 \
NONOBENCH_LOCAL_MODEL=Qwen3-32B \
bun run bench:local --model Qwen3-32B
Do a pilot first. The next command runs two 5x5 puzzles into a scratch database:
NONOBENCH_LOCAL_BASE_URL=http://192.168.1.20:8000/v1 \
NONOBENCH_LOCAL_MODEL=Qwen3-32B \
NONOBENCH_DB=local-pilot.db \
bun run bench --model Qwen3-32B --sizes 5x5 --limit 2
Make sure that the plan lists your model as [local, text].
Environment variables for a local model
| Variable | Meaning |
|---|---|
NONOBENCH_LOCAL_BASE_URL | The /v1 endpoint of the server. Ollama uses port 11434. |
NONOBENCH_LOCAL_MODEL | The model id that the server returns from GET /v1/models. |
NONOBENCH_LOCAL_NAME | The display name and the value for --model. Defaults to the model id. Must not match a cloud model name. |
NONOBENCH_LOCAL_EFFORT | The reasoning effort your server is configured with (none, low, xhigh, …). Label only — the runner never sends reasoning settings to a local server. Defaults to none. |
NONOBENCH_LOCAL_API_KEY | The bearer token, when the server requires one. Defaults to local. |
The runner adds the local model to the plan only when you set both NONOBENCH_LOCAL_BASE_URL and NONOBENCH_LOCAL_MODEL.
Notes on local runs
- Local runs cost $0. Keep them out of the published dataset: run them with
bun run bench:local, which isbun run benchwithNONOBENCH_DB=local-results.db.bench/results.dbis the shared dataset, and the export-contract test checks it against the committed exports. The runner refuses to write a local model to the shared database. - Before starting, the runner checks the model
idagainst the server'sGET /v1/modelslist. A typo fails with the models the server does serve, and an unreachable server fails before any run is recorded. - Text mode is the default for a local model. Many local servers accept a JSON schema request and then ignore it. If your server enforces the schema, set
NONOBENCH_OUTPUT_MODE=json_schema. - Start with
--parallel 1. One GPU serves fewer requests at the same time than a cloud provider does. - The 20x20 tier asks for 128,000 output tokens. If your server has a smaller context, run only the core sizes.
- The runner does not apply the reasoning-token circuit breaker to a local model. Some local servers report zero reasoning tokens for every run.
- Set the same variables in
bench/.envto avoid the prefix on every command. - A bench run registers the local model in
bench/local-models.json(name, server model id, family, server URL), so a laterbun run exportand the visualizer pick it up without any env vars. To register a model you already benched, add one entry there:"<name>": { "baseURL": "http://host:port/v1", "modelId": "<server model id>", "family": "<name>" }. OmitmodelIdwhen the name is the server's id.
4. Viewing Results
cd visualizer
bun install
bun run dev
Then open http://localhost:3000 to view the interactive dashboard.
Your own runs live in a separate database. Serve them with:
cd visualizer
bun run dev:local
dev:local exports bench/local-results.db into the dashboard's data files, starts the dev server, and restores the committed files when you stop it. While it runs, the dashboard shows only the models in that database — your local runs, not the published leaderboard. Set NONOBENCH_LOCAL_DB to serve another database.
While that server runs, bun test in bench fails one test: the export-contract test compares the dashboard's data files against bench/results.db, and dev:local has put your runs in them. Stop the server, let the script restore the committed files, and the suite passes again.
5. Deploying
The visualizer is a Cloudflare Worker (visualizer/wrangler.jsonc). Pages are prerendered at build time and served as static assets; the Worker serves the API, the MCP server and the agent routes, reading the exported JSON through its assets binding. Merging to main deploys production through Cloudflare Workers Builds. From visualizer/:
bun run deploy:betabuilds the checkout and deploys it to beta.nonobench.com, a separate Worker that sendsX-Robots-Tag: noindex.bun run deploy:verify <sha>waits until the live Worker runs a build containing that commit: production for commits onmain, the beta for branch commits.
Agent Access
nonobench.com exposes the benchmark data to agents, with no authentication:
-
REST API under
/api/v1(leaderboard, models, puzzles, a solution checker, individual runs). The spec is at/api/openapi.json, and/.well-known/api-catalog(RFC 9727) points to it. -
MCP server at
/mcp(stateless Streamable HTTP, MCP 2026-07-28 with 2025 client compatibility), described by/.well-known/mcp/server-card.json. Add it to a client withclaude mcp add --transport http nonobench https://www.nonobench.com/mcp. Browser requests may use the site's origins or HTTP localhost/127.0.0.1 origins. It's listed in the Claude Connectors Directory, the MCP Registry (io.github.mauricekleine/nonobench), Smithery, Glama and mcpservers.org. -
WebMCP tools registered in the browser via
navigator.modelContext. -
Markdown:
/and/puzzlesreturn markdown when requested withAccept: text/markdown./llms.txtgives an overview. -
Discovery:
robots.txt(with Content Signals),sitemap.xml,Linkheaders on the homepage, an agent skill at/.well-known/agent-skills/index.json, and an ARD manifest at/.well-known/ai-catalog.json.
All of it is read from the same exported files as the dashboard (visualizer/app/results.json and visualizer/public/results-raw.json), so bun run export updates it too.
Grading
Each model receives the same system prompt and the puzzle's row and column clues. Standard answers are the grid as one string of 1s and 0s. Hard mode answers are one row per line, because at 400 cells most models miscount a single string (see LEARNINGS.md). An answer is correct when it satisfies every row and column clue.
Ten of the 30 puzzles (one 5x5, four 10x10, five 15x15) have more than one valid solution, so answers are checked against the clues rather than compared with the stored solution. Correctness is derived from the stored raw outputs at export time; the database is never rewritten. The puzzle test suite pins which puzzles are ambiguous, and any new puzzle must have a unique solution.
Puzzle Data
The core tier has 30 puzzles (10 each of 5x5, 10x10 and 15x15), defined in visualizer/components/puzzles/ and shared by the runner and the dashboard. They were sourced from nono-dataset. Hard mode has 10 generated 20x20 puzzles. A puzzle's ID is a hash of its solution, so changing a puzzle's solution creates a new puzzle.
Tiers and generation
Default benchmark runs cover the three core sizes (Standard). Use --sizes 20x20 with a model selection to run Hard mode; comma-separated sizes also work. The runner's plan reports missing 20x20 work separately. Headline overall accuracy and best-variant selection use Standard runs only; Hard mode has its own results. Hard-mode requests get a 128,000-token answer budget, capped at the endpoint's maximum (bench/max-output-tokens.json).
From bench/, bun run generate-puzzles recreates the 20x20 set with a fixed seed. It fills grids at random (no pictures, so a model can't guess the image), keeps grids with at least three blocks per line and little mirror symmetry, and checks uniqueness with an exact solver (NONOGRAM_SOLVER). The set mixes five puzzles that row-and-column propagation solves with five where it stalls with 20–200 cells left. bun test verifies their clues, uniqueness flags and line solvability; the original ambiguity list remains pinned.
Configuration
Edit bench/constants.ts to configure:
MODELS- Array of model configurations (OpenRouter model ID, display name, reasoning settings)MAX_PARALLEL_RUNS_PER_MODEL- Concurrent puzzle runs per model (default: 10)REQUEST_TIMEOUT_MS- Per-request timeout; a timed-out request is stored as a failed run (default: 30 minutes)
NONOBENCH_DB, NONOBENCH_RESULTS_JSON and NONOBENCH_RESULTS_RAW_JSON override the database and export paths, which is useful for testing against a copy.
Tech Stack
Benchmark Runner
- Bun - JavaScript runtime and SQLite
- AI SDK - Unified LLM interface
- OpenRouter - LLM API gateway
- TypeScript
Visualizer
- TanStack Start - React framework, prerendered pages
- Cloudflare Workers - hosting, with static assets
- React 19 - UI library
- Tailwind CSS 4 - Styling
- shadcn/ui - Component library
- Recharts - Charts
- Zustand - State management
Contributing
Contributions are welcome! Feel free to:
- Add support for new LLM models
- Improve the benchmark methodology
- Enhance the visualization dashboard
License
MIT
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Git
Freeby Modelcontextprotocol · Developer Tools
Read, search, and manipulate Git repositories programmatically
Paperclip
Freeby Paperclipai · Developer Tools
Trending hip-hop artist momentum scores across four cultural dimensions.
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
MCP Marketplace
Freeby mcp-marketplace · Developer Tools
Search and install MCP servers from inside your AI client.
