Back to Browse

Pdf Triage MCP Server

Developer ToolsUse Caution3.0MCP RegistryLocal
Free

Server data from the Official MCP Registry

Read local PDFs without uploading. Classifies first, flags untrustworthy text, bounds output.

About

Read local PDFs without uploading. Classifies first, flags untrustworthy text, bounds output.

Security Report

3.0
Use Caution3.0High Risk

Valid MCP server (1 strong, 1 medium validity signals). 5 known CVEs in dependencies (2 critical, 3 high severity) Package registry verified. Imported from the Official MCP Registry.

5 files analyzed · 6 issues found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

HTTP Network Access

Connects to external APIs or services over the internet.

env_vars

Check that this permission is expected for this type of plugin.

Unverified package source

We couldn't verify that the installable package matches the reviewed source code. Proceed with caution.

What You'll Need

Set these up before or after installing:

Minimum severity written to stderr. Defaults to info.Optional

Environment variable: PDF_TRIAGE_LOG_LEVEL

Default truncation ceiling for tool responses, in characters. Guards the caller's context window. Defaults to 40000.Optional

Environment variable: PDF_TRIAGE_MAX_CHARS

Largest PDF the server will read, in bytes. The parser buffers whole files rather than streaming. Defaults to 104857600 (100 MB).Optional

Environment variable: PDF_TRIAGE_MAX_FILE_BYTES

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-vishalmeena2211-pdf-triage-mcp": {
      "env": {
        "PDF_TRIAGE_LOG_LEVEL": "your-pdf-triage-log-level-here",
        "PDF_TRIAGE_MAX_CHARS": "your-pdf-triage-max-chars-here",
        "PDF_TRIAGE_MAX_FILE_BYTES": "your-pdf-triage-max-file-bytes-here"
      },
      "args": [
        "-y",
        "pdf-triage-mcp"
      ],
      "command": "npx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

pdf-triage-mcp

npm CI MCP Registry License: MIT

An MCP server that lets any AI tool read local PDFs — without uploading them, without an OCR bill, and without silently handing back garbage.

Built on @firecrawl/pdf-inspector (Rust, no ML models, no external services).

┌─ pdf_classify ──→  text_based · 0.98 · 12 pages · 0 need OCR   ~20ms
├─ pdf_search  ──→  "invoice total" found on p4, p9              ~150ms
├─ pdf_extract ──→  clean Markdown, truncated to your budget     ~150ms
└─ pdf_tables  ──→  just the pipe tables, no prose               ~150ms

Table of contents


Why this exists

Most PDF tooling has the same failure mode: it returns confident text regardless of whether extraction actually worked. Broken CID fonts, substitution-cipher encodings, scanned pages with no text layer — you get plausible-looking output and find out downstream, if at all.

pdf-inspector is unusually good at knowing when it failed. It emits U+FFFD rather than guessing at an unmapped CID, runs substitution-cipher detection over its own output, and reclassifies a document as scanned when extracted text drops below 50% alphanumeric. But it stops at reporting those findings on a result object — and most wrappers throw them away.

This server acts on them. Every response carries the trust signals, above the content, where the model reads them first:

> [!WARNING] ENCODING ISSUES DETECTED. The text layer decoded to suspicious
> output — typically a garbled CID font or a substitution-cipher encoding
> where letter frequencies match natural language but the letters themselves
> are wrong. Treat all extracted text here as unreliable and prefer OCR.

---

# Quarterly Report
...

Three design rules follow:

  1. Classify before extracting. pdf_classify costs ~20ms and tells you whether extraction is worth attempting at all.
  2. Bound every output. A 300-page PDF is easily 500k tokens. Everything truncates by default and tells you how to page through instead.
  3. Confine every path. A model that has just read an untrusted document must not be talkable into reading ~/.ssh/id_rsa. Enforced in code, not left to the model's judgement.

Install

Nothing to install. Every config below runs the published package straight from npm:

npx -y pdf-triage-mcp --root /path/to/your/documents

Your MCP client runs that for you — you only need to paste the config. Confirm it works first:

npx -y pdf-triage-mcp --version

Requires Node 20+. Available on npm as pdf-triage-mcp and in the MCP Registry as io.github.vishalmeena2211/pdf-triage-mcp.

git clone https://github.com/vishalmeena2211/pdf-triage-mcp.git
cd pdf-triage-mcp
npm install
npm run build
node dist/index.js --root ~/Documents

Then substitute "command": "node", "args": ["/absolute/path/to/dist/index.js", ...] for the npx invocation in any config below.


Connect it to your AI tool

Every config below is complete as written except for one value:

  • /Users/me/Documents — replace with the directory the server may read. This is the only thing you must change.

Use an absolute path; ~ is not expanded by most clients. Repeat --root for multiple directories.

PATH gotcha, applies to every GUI client below. Desktop apps launch servers with a minimal environment, so bare npx often fails to resolve even though it works in your terminal. If the server won't start, substitute the absolute path — find it with which npx (commonly /opt/homebrew/bin/npx on Apple Silicon, /usr/local/bin/npx on Intel macOS).

Why -y? It skips npx's install confirmation prompt. Without it, a first run can hang waiting for input that an MCP client cannot provide — the server appears to start and then silently times out.

Config file

OSPath
macOS~/Library/Application Support/Claude/claude_desktop_config.json
Windows%APPDATA%\Claude\claude_desktop_config.json
Linux~/.config/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}

Verify: Fully quit and relaunch Claude Desktop (not just close the window). A tools icon appears near the chat input — click it and confirm the four pdf_* tools are listed.

Logs: ~/Library/Logs/Claude/mcp*.log (macOS), %APPDATA%\Claude\logs\mcp*.log (Windows).

Docs

CLI — the easiest route. The -- separator is mandatory; everything after it is the server command.

# Just you, this project (default)
claude mcp add pdf-triage -- npx -y pdf-triage-mcp --root /Users/me/Documents

# Just you, every project
claude mcp add --scope user pdf-triage -- npx -y pdf-triage-mcp --root /Users/me/Documents

# Shared with your team, writes .mcp.json to the repo
claude mcp add --scope project pdf-triage -- npx -y pdf-triage-mcp --root /Users/me/Documents

Or edit .mcp.json at the project root directly:

{
  "mcpServers": {
    "pdf-triage": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}
ScopeStored inShared
local (default)~/.claude.json, under this projectNo
user~/.claude.json, top levelNo
project.mcp.json in repo rootYes, via git

Verify: claude mcp list → look for ✔ Connected. Project-scoped servers need approval on first use — run /mcp inside a session.

Docs

Config file: .cursor/mcp.json (project) or ~/.cursor/mcp.json (global). Project wins on conflict.

{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}

Verify: Cursor hot-reloads — no restart. Open Cursor Settings → Tools & MCP and look for a green dot next to pdf-triage.

Docs

Config file: ~/.codeium/windsurf/mcp_config.json (macOS/Linux), %USERPROFILE%\.codeium\windsurf\mcp_config.json (Windows).

Not created on first launch — create it yourself if missing.

{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}

Verify: Windsurf watches the file and hot-reloads on save. Tools appear in Cascade on the next chat session.

Docs

The key is servers, not mcpServers. This is the most common mistake when copying a config from Claude Desktop.

Config file: .vscode/mcp.json (workspace), or Command Palette → MCP: Open User Configuration (global).

{
  "servers": {
    "pdf-triage": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}

CLI alternative:

code --add-mcp '{"name":"pdf-triage","command":"npx","args":["-y","pdf-triage-mcp","--root","/Users/me/Documents"]}'

Verify: MCP tools only work in Agent mode — switch from Ask/Edit to Agent in Copilot Chat, then click Configure Tools and confirm the pdf_* tools appear. Restart VS Code after first adding the file.

Docs

The key is context_servers, not mcpServers, and command is a nested object rather than a string.

Config file: ~/.config/zed/settings.json (macOS/Linux), %APPDATA%\Zed\settings.json (Windows). Command Palette → zed: open settings.

{
  "context_servers": {
    "pdf-triage": {
      "source": "custom",
      "command": {
        "path": "npx",
        "args": [
          "-y", "pdf-triage-mcp",
          "--root", "/Users/me/Documents"
        ],
        "env": {}
      }
    }
  }
}

If your Zed version rejects that, it predates the nested form — try command, args and env flat at the top level of the server object instead.

Verify: Agent Panel (Cmd+Shift+A) → gear icon → MCP Servers. Green dot means connected.

Docs

Config file — separate from VS Code's own:

OSPath
macOS~/Library/Application Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json
Windows%APPDATA%\Code\User\globalStorage\saoudrizwan.claude-dev\settings\cline_mcp_settings.json
Linux~/.config/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json
{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ],
      "disabled": false,
      "autoApprove": ["pdf_classify", "pdf_search"]
    }
  }
}

autoApprove runs the listed read-only tools without a confirmation prompt.

Easier route: Cline panel → MCP servers icon → Edit MCP Settings opens this file directly.

Verify: Panel refreshes automatically; green dot next to the server.

Docs

Config file: ~/.continue/config.yaml (global) or .continue/config.yaml (project). YAML is current; config.json is deprecated.

mcpServers here is a list, not an object — and YAML needs spaces, never tabs.

mcpServers:
  - name: pdf-triage
    command: npx
    args:
      - -y
      - pdf-triage-mcp
      - --root
      - /Users/me/Documents

Verify: Reloads automatically on save. Switch Continue to Agent mode — MCP tools are unavailable in other modes.

Docs

CLI:

gemini mcp add pdf-triage npx -y pdf-triage-mcp --root /Users/me/Documents

# global instead of project-scoped
gemini mcp add --scope user pdf-triage npx -y pdf-triage-mcp --root /Users/me/Documents

Or edit ~/.gemini/settings.json (global) / .gemini/settings.json (project):

{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ],
      "timeout": 30000,
      "trust": false
    }
  }
}

Verify: Run /mcp inside a gemini session — servers show CONNECTED with their tool list. Or gemini mcp list from the shell.

Docs

Config file: ~/.codex/config.toml (global) or .codex/config.toml (project).

TOML, and the key is mcp_servers — snake_case, never mcpServers.

[mcp_servers.pdf-triage]
command = "npx"
args = [
  "-y", "pdf-triage-mcp",
  "--root", "/Users/me/Documents"
]
startup_timeout_sec = 20
tool_timeout_sec = 60

Verify: codex doctor --json validates the config syntax. Note that it validates syntax only — it does not confirm the server actually spawned.

Known upstream issue: several Codex CLI versions have a bug where stdio servers validate cleanly but silently fail to start, showing Tools: none in the TUI (#3441, #26810). That is a Codex runtime bug, not a config error.

Docs

AI Assistant — configured through the IDE, no file to edit:

  1. Settings → Tools → AI Assistant → Model Context Protocol (MCP)
  2. Add → transport STDIO
  3. Paste:
{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}
  1. OK → Apply

Junie uses a file instead — ~/.junie/mcp/mcp.json (global) or .junie/mcp/mcp.json (project), same JSON shape.

Verify: Check the Status column in the MCP settings panel; click it to list the server's tools.

Docs

Config file: ~/.lmstudio/mcp.json (macOS/Linux), %USERPROFILE%\.lmstudio\mcp.json (Windows).

Easier via the app: right sidebar → Program tab → Install → Edit mcp.json.

{
  "mcpServers": {
    "pdf-triage": {
      "command": "npx",
      "args": [
        "-y", "pdf-triage-mcp",
        "--root", "/Users/me/Documents"
      ]
    }
  }
}

Verify: Auto-reloads on save; tools appear in the Program panel. LM Studio shows a confirmation dialog the first time a model calls a tool.

Docs

Cheat sheet

ClientFileTop-level keyRestart?
Claude Desktopclaude_desktop_config.jsonmcpServersFull quit
Claude Code.mcp.json / CLImcpServersNo
Cursor.cursor/mcp.jsonmcpServersNo
Windsurf~/.codeium/windsurf/mcp_config.jsonmcpServersNo
VS Code Copilot.vscode/mcp.jsonserversFirst time
Zed~/.config/zed/settings.jsoncontext_serversNo
Clinecline_mcp_settings.jsonmcpServersNo
Continue.dev~/.continue/config.yamlmcpServers (list)No
Gemini CLI~/.gemini/settings.jsonmcpServersNo
Codex CLI~/.codex/config.toml[mcp_servers.*]N/A
JetBrainsIDE settings UImcpServersNo
LM Studio~/.lmstudio/mcp.jsonmcpServersNo

The three that differ: VS Code (servers), Zed (context_servers + nested command), Codex (TOML mcp_servers). Everything else takes the Claude Desktop format verbatim.


Tools

ToolCostPurpose
pdf_classify~20msType, confidence, page count, exact pages needing OCR. Call this first.
pdf_extract~150msPDF → Markdown. Truncates by default; slice with pages.
pdf_search~150msLocate text, return page-attributed snippets. Cheapest way into a long document.
pdf_tables~150msTables only, as Markdown pipe tables.

Full parameter reference: docs/TOOLS.md.

The intended flow on an unfamiliar document:

pdf_classify → is it text_based with no warnings?
  ├─ yes → pdf_search to locate → pdf_extract with `pages`
  └─ no  → stop; route to OCR

Configuration

pdf-triage-mcp [options]

  -r, --root <dir>          Directory the server may read. Repeatable. Default: cwd.
      --max-chars <n>       Default truncation ceiling. Default: 40000. Max: 200000.
      --max-file-bytes <n>  Largest PDF to read. Default: 104857600 (100 MB).
      --log-level <level>   debug | info | warn | error | silent. Default: info.
  -h, --help                Show usage.
  -v, --version             Print version.

Environment equivalents: PDF_TRIAGE_ROOTS (separated by the platform PATH delimiter — : on macOS/Linux, ; on Windows), PDF_TRIAGE_MAX_CHARS, PDF_TRIAGE_MAX_FILE_BYTES, PDF_TRIAGE_LOG_LEVEL. Flags win over environment.

Roots are a security boundary, not a convenience. Grant the narrowest directory that works. Paths are resolved through symlinks before checking, so a link inside a root pointing outside it is rejected rather than followed.


Engine fallback

Upstream ships prebuilt native binaries for exactly three targets: linux-x64-gnu, darwin-arm64, win32-x64-msvc. No musl build, no Linux ARM64 build (upstream #216) — so it fails to load on Alpine containers, Graviton instances, and most edge runtimes.

This server prefers native and falls back to WASM, which runs anywhere. Capability differences are surfaced, never faked:

NativeWASM
Classify / extractYesYes
Per-page extractionYesNo — throws, and pdf_search reports its matches are unattributed
pages selectionYesNo — ignored, and the response says so

Check which engine you got: pdf_classify reports it, and the server logs engine selected at startup.


Known limitations

Inherited from upstream. Worth reading before you trust output:

  • RTL scripts are broken. Arabic and Hebrew return in visual order, reversed and unusable, while being reported as text_based with high confidence (#212). This server detects and escalates it — the one upstream failure mode we actively guard.
  • No xref recovery. Malformed PDFs that pypdf and pdfium silently repair will throw (#228).
  • Japanese CIDFontType0 (CFF) subset fonts can decode to unrelated glyphs (#208).
  • Multi-column reading order may emit in raster order on some layouts despite columns being detected (#219).
  • Images are not PDFs. A scanned JPEG has no text layer; the server rejects non-PDF input rather than pretending otherwise.

Development

npm run typecheck    # tsc --noEmit, maximum strictness
npm test             # vitest, 108 tests
npm run test:coverage
npm run build
npm run dev          # tsx, no build step

The TypeScript config runs every strictness flag including exactOptionalPropertyTypes and noUncheckedIndexedAccess. Upstream responses are validated with Zod at the boundary rather than cast — see docs/ARCHITECTURE.md for why.

Debug a client connection:

node dist/index.js --root ~/Documents --log-level debug

Troubleshooting: docs/TROUBLESHOOTING.md.


Roadmap

  • pdf_regions — bbox-scoped extraction for hybrid model pipelines
  • Positioned-item tool exposing {page, bbox} for visual citation UX
  • Integration tests asserting native and WASM produce identical normalized shapes
  • Optional OCR adapter interface, closing the routing loop end to end

License

MIT

Reviews

No reviews yet

Be the first to review this server!