Back to Browse

TensorCAD MCP Server

Developer ToolsLow Risk9.7MCP RegistryLocal
Free

Server data from the Official MCP Registry

Design transformer LLM architectures and report their parameters, FLOPs, memory and cost

About

Design transformer LLM architectures and report their parameters, FLOPs, memory and cost

Security Report

9.7
Low Risk9.7Low Risk

Valid MCP server (2 strong, 1 medium validity signals). No known CVEs in dependencies. ⚠️ Package registry links to a different repository than scanned source. Imported from the Official MCP Registry. 1 finding(s) downgraded by scanner intelligence.

13 files analyzed · 1 issue found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

What You'll Need

Set these up before or after installing:

Directory the server scans for .tensorcad.json files and resolves relative paths against. Defaults to the process working directory.Optional

Environment variable: TENSORCAD_ROOT

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-filip-pajalic-tensorcad": {
      "env": {
        "TENSORCAD_ROOT": "your-tensorcad-root-here"
      },
      "args": [
        "-y",
        "@tensor-cad/mcp"
      ],
      "command": "npx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

TensorCAD


TensorCAD treats a neural network the way an EDA tool treats a circuit. Blocks are symbols with typed pins. Tensors are nets. Shapes are checked by a real algebra rather than by running the thing. A design-rule check tells you the model will not fit on your GPUs before you rent them. And the drawing is not a picture of the model — it is the model, and PyTorch falls out of it.

It started as a tool for language models. It now draws vision transformers and convolutional classifiers too, because the same machinery turned out to work.

Every picture here is regenerated by bun run scripts/screenshots.ts from a running dev server, so it is what the tool looks like rather than what it looked like once. bun run scripts/export-svg.ts writes the sheet itself as a vector — one is in docs/images — which is also what File > Export the sheet as SVG does from the editor.

It runs in a browser: tensorcad.dev. There is no server behind it — the engine is the same WebAssembly module the command line and the MCP server load, so every number on the screen is computed in the tab, and this build has nowhere to send a design even if it wanted to: it registers no storage provider, so File > Save a copy writes to your disk and that is the only copy there is.

A separate deployment at app.tensorcad.dev is this same editor with an account attached, where designs are saved and can be shared. The first claim holds there too — the analysis still runs in the tab — but the second does not, which is why they are two sentences and two addresses rather than one of each.

The documentation is at docs.tensorcad.dev.

Quick start

Needs Bun and Go 1.25 or later. The analysis engine is Go compiled to WebAssembly, which is the one build step:

git clone https://github.com/Filip-Pajalic/TensorCAD
cd TensorCAD
bun install
bun run build:wasm

Every number in the project comes out of one command:

bun run scripts/report.ts
preset              calculated     published   delta
------------------------------------------------------------
gpt2-small              124.4M        124.4M   exact
llama-3-8b               8.03B         8.03B   exact
deepseek-v3            671.03B       671.00B   0.0039%
ijepa-vit-h14            1.28B         1.28B   exact
alexnet                  61.1M         61.1M   exact

Generate a model and check it against real PyTorch:

bun run scripts/codegen-demo.ts llama-3-8b      # writes out/llama-3-8b/model.py
python -m tensorcad_runtime verify out/llama-3-8b/model.py

Open the editor:

bun run --cwd packages/ui dev      # browser
cd desktop && wails3 task build    # desktop app (Wails v3 + Go)

Both load the same engine. So do the command line and the MCP server, which is the point: the numbers cannot depend on where you asked for them.

What it does

Schematic capture. Blocks, typed pins, orthogonal wire routing, four-sided pin anchors, junction dots on branching nets, hollow circles on unconnected pins. Containers unfold in place so a 32-layer stack reads as one frame with a 32× bracket, the way published architecture figures draw it.

A feature timeline. Every edit is kept as an operation with its arguments, not as a snapshot, so a step can be taken out of the middle and everything after it replays on top of what is left. Suppress the step that widened the model and the rename you did afterwards survives. A step that cannot replay — because you suppressed the one that added the block it wired — says what it could not find rather than being dropped in silence.

Tensors you can point at. A wire is a tensor, and clicking one says what it carries: its shape, its dtype, the block that made it, every block that reads it, and its share of the activation memory. Every segment of the same net lights with it. A block that fans out — Nemotron-H's split holds 290 MiB across three output pins, 128, 160 and 2 — is where that matters: its own number answers neither which of them is the big one nor what dropping one would save.

A real shape algebra. Every tensor dimension is a multivariate polynomial with exact rational coefficients over named symbols. B and T stay indeterminate all the way through, so a mismatch is a genuine polynomial difference rather than two numbers that happened not to match. Splits carry divisibility obligations instead of silently rounding.

Design-rule checks. Eighteen rules: head divisibility, vocabulary padding, RoPE dimension parity, interface breakage, whether the design fits the GPUs you selected under the sharding plan you chose. The DRC panel correctly refuses Llama-3-8B at 90.16 GiB/GPU against an H100's 80.

Quantitative analysis. Parameters, FLOPs, activation memory (Megatron formulas), KV cache, ZeRO/FSDP/TP/PP sharding, roofline throughput, Chinchilla budgets. Nothing in the UI computes its own numbers; one validate() call per document and operating point feeds every panel.

PyTorch generation. generateTorch(doc) emits a runnable model with an init_weights() method — because nn.Embedding defaults to a unit normal, and that is the difference between a next-token loss of 466 and 10.94 against the ln(50257) = 10.82 baseline.

A 3D volume view. Every tensor as a plate, sized by its real dimensions, with flow ribbons between them. Ported from Brendan Bycroft's LLM visualisation.

An MCP server. So an agent can design, validate, analyse and generate without a human in the loop.

How it is kept honest

This is the part worth reading, and the reason to trust the numbers.

Twenty presets are the regression suite. Each one carries the parameter count its authors published, and the tests assert the analysis reproduces it. Seventeen match to the parameter; the other three are checked against rounded vendor figures with an explicit tolerance.

Every preset is instantiated in real PyTorch. python -m tensorcad_runtime verify builds the generated model on the meta device and reports its true parameter count, module by module, up to DeepSeek-V3 at 671,026,419,200.

FLOPs are checked against a profiler. For AlexNet the agreement is exact — 1,428,376,960 per image, ratio 1.000000 against torch.utils.flop_counter. For GPT-2 small the profiler says 251.78 MFLOP/token and the analysis says 249.42, and the whole difference is the causal mask: a profiler counts the attention operator as if nothing were masked. flops.fwdTotalUnmasked reproduces the profiler exactly; flops.fwdTotal is what a fused causal kernel actually does. A test pins both numbers.

The Go port is proven against the TypeScript it replaces. The engine is migrating to Go; the TypeScript writes golden files for all twenty presets and the Go tests must reproduce them exactly — including the evaluation order of the symbol table, the printed form of every polynomial, and the text of every error.

Architectures it draws

LanguageGPT-2 (small→XL), nanoGPT, Llama 2/3/3.1, Mistral, Qwen 2.5/3, Gemma 2, Mixtral, DeepSeek-V3, Nemotron-H
VisionI-JEPA ViT-H/14 — bidirectional attention, three towers including the EMA target encoder
ConvolutionalAlexNet — B C H W tensors, spatial downsampling, 96% of its weights in the classifier

Mechanisms covered: GQA/MQA/MHA, multi-head latent attention, SwiGLU/GeGLU, RMSNorm/LayerNorm, RoPE with scaling, mixture-of-experts with shared experts and routing bias, Mamba-2 state-space layers, sliding-window attention, QK-norm, post-norm, tied embeddings, hybrid stacks.

Repository shape

packages/core-go/   the engine — Go, no dependencies outside the standard library
packages/engine/    the engine compiled to WebAssembly, and its TypeScript client
packages/ui/        React + React Flow editor, 2D sheet and 3D volume view
packages/cli/       command line: validate, analyze, show, diff, codegen
packages/mcp/       MCP server
desktop/            Wails v3 + Go desktop application
python/             the only Python: verifies generated models against PyTorch
docs/               tutorials, how-to guides, reference and explanation

One engine, everywhere. The editor, the command line, the MCP server and the desktop shell all load the same WebAssembly module and ask it the same questions, so an answer cannot depend on where it was asked. What it is held to is packages/core-go/testdata: every preset's symbol table, inferred shapes, full analysis, design-rule findings and generated PyTorch, byte for byte, checked both against the Go source and against the compiled module. Those files began as the answers of the TypeScript this was ported from, which has since been deleted.

For agents

This repository is written to be worked on by coding agents as well as people.

  • CLAUDE.md is the entry point: commands, layout, the eight invariants, and what to do when adding a block. Read it first.
  • Invariants are load-bearing. The document is the source of truth; formulas live only on primitives; containers carry two multipliers; activation memory is attributed to tensors rather than blocks; B and T are reserved. Break one and the tests will tell you, but the design rules will not.
  • Adding a block needs parameter specs, port patterns, docs.summary and docs.formula with a source link, and — for a primitive — paramCount, flops, retains and stateBytes. Then a preset that uses it with a published figure, or a test pinning the arithmetic. Run bun run scripts/report.ts before and after.
  • The MCP server exposes the whole engine as tools. Point your agent at .mcp.json.
  • Start it with TENSORCAD_BRIDGE=1 and the editor attaches to it. The agent's edits appear on the canvas as it makes them and land on the undo stack, so a person watching can take one back; what that person does comes back the other way. There is one document, not two — the editor's edits go through the same revision-checked apply a tool call does. It binds 127.0.0.1, refuses a foreign Origin, and opens no port at all without the variable. See Drive TensorCAD from an agent.

Built on

TensorCAD borrows from work that deserves naming:

  • llm-viz by Brendan Bycroft (MIT) — the 3D volume view is a port of its layout and arrow rendering. The residual pathway down the centre, weights to either side, blocks wrapping into columns, the ribbon arrows with lines down their edges: the arrangement is his.
  • KiCad — the interaction model. Pick apertures, net highlighting, junction dots, dangling-pin marks and the selection semantics are taken from eeschema's behaviour and documentation.
  • nanoGPT by Andrej Karpathy (MIT) — a preset, and the reference for GPT-2-shaped arithmetic.
  • I-JEPA by Meta AI — the vision preset is built from its published configuration.
  • torchvision (BSD-3) — the AlexNet definition and its parameter count.
  • React Flow, ELK, Three.js, Base UI, Wails — the editor stands on these.

Formulas are sourced individually in docs/reference/analysis-math.md, including two figures the original research got wrong that the implementation corrects.

Status

Working and useful, with rough edges. The Go migration is at stage 2 of 10. See ROADMAP.md for what is known to be missing — linear-attention blocks, multi-token prediction, and Gemma's alternating local/global attention, which the importer warns about rather than approximating.

Documentation

docs/, organised by Diátaxis: tutorials to learn from, how-to guides to work from, reference to look things up in, and explanation for why any of it is the way it is.

License

MIT. See LICENSE.md, which also carries the notices for the MIT-licensed work this project ports.

Reviews

No reviews yet

Be the first to review this server!