Server data from the Official MCP Registry
URL to LLM-ready markdown — plus per-page category, page_structure, and query-driven highlights.
About
URL to LLM-ready markdown — plus per-page category, page_structure, and query-driven highlights.
Security Report
A well-structured MCP server for web content extraction with solid security practices. Authentication is properly required via API key (OCTEN_API_KEY environment variable), input validation is comprehensive, and the codebase is clean with no malicious patterns or dangerous operations. Minor code quality observations around error handling and logging do not significantly impact the overall security posture. Supply chain analysis found 2 known vulnerabilities in dependencies (0 critical, 2 high severity). Package verification found 1 issue.
4 files analyzed · 7 issues found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Required
This plugin requests these system permissions. Most are normal for its category.
What You'll Need
Set these up before or after installing:
Environment variable: OCTEN_API_KEY
Environment variable: OCTEN_API_URL
How to Install
Add this to your MCP configuration file:
{
"mcpServers": {
"io-github-octen-team-octen-mcp": {
"env": {
"OCTEN_API_KEY": "your-octen-api-key-here",
"OCTEN_API_URL": "your-octen-api-url-here"
},
"args": [
"-y",
"octen-mcp"
],
"command": "npx"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
octen-mcp
MCP server for Octen. Plug it into Claude, Cursor, VS Code, Windsurf, or any MCP client to give your agent live web search and URL extraction.
Core capabilities:
search/news_search: search the live web with domain, text, language, and time filters.broad_search: decompose a query into multiple sub-queries, search them concurrently, and return results grouped per sub-query for broad coverage.extract: turn one or more URLs into clean, LLM-ready content.image_search(In Beta — contact us for beta access): search the web for images by text query, optionally with a reference image.video_search(In Beta — contact us for beta access): search the web for videos by text query.
What makes Octen useful for agents is that extract returns more than page text. Each successful result also includes:
category: what the page is aboutpage_structure: what kind of page it ishighlights: ranked snippets when you pass aquery
That lets an agent skip login walls, nav pages, and off-topic URLs before spending tokens on the full body.
Why Octen MCP
Fast
Web search averages 62ms. Fast enough for multi-step MCP workflows.
Accurate
Powered by SOTA text and VL embedding models. Better sources, fewer hallucinations.
Fresh
Live web data with minute-level updates. Useful for news, prices, and fast-moving pages.
Efficient
Clean highlights, optional full_content, and page labels keep model context relevant.
Quick start
You need an OCTEN_API_KEY from octen.ai.
Node compatibility: 0.4.2 and later run on every supported Node (>= 18.17), including Node 26+. Versions 0.4.1 and below fail every call on hosts whose embedded undici is v8+ (Node 26 and later) with
Network error … code=UND_ERR_INVALID_ARG cause=invalid onError method— if you see that error, upgrade the package (or run on Node <= 24).
For most MCP clients, the config is:
{
"mcpServers": {
"octen": {
"command": "npx",
"args": ["-y", "octen-mcp"],
"env": {
"OCTEN_API_KEY": "your-key-here"
}
}
}
}
Install command by client
| Agent | One-line install |
|---|---|
| Claude Code | claude mcp add --scope user octen -e OCTEN_API_KEY=your-key-here -- npx -y octen-mcp |
| Codex | codex mcp add octen --env OCTEN_API_KEY=your-key-here -- npx -y octen-mcp |
| Gemini CLI | gemini mcp add octen -e OCTEN_API_KEY=your-key-here -- npx -y octen-mcp |
| VS Code | code --add-mcp '{"name":"octen","command":"npx","args":["-y","octen-mcp"],"env":{"OCTEN_API_KEY":"your-key-here"}}' (or click a badge above) |
| Cursor | Add to Cursor (then edit the key), or use the JSON above in ~/.cursor/mcp.json |
| Claude Desktop | No CLI — add the JSON above to the config file (see below) |
Config file locations
For clients without a CLI installer, drop the JSON config above into:
- Claude Desktop:
~/Library/Application\ Support/Claude/claude_desktop_config.json - Cursor:
~/.cursor/mcp.json - VS Code workspace:
.vscode/mcp.json(useserversinstead ofmcpServers) - Windsurf / Cline / other clients: paste it into that client's MCP settings
Tools
| Tool | What it does | Best for |
|---|---|---|
search | Search the live web with domain, text, language (ISO 639-1), time, and content controls | a single focused web search |
news_search | Same engine as search, fixed to news | current events and timely reporting |
broad_search | Decompose a query into up to max_queries sub-queries, search concurrently, return grouped results (same per-sub-query options as search, including the language filter) | research-style, multi-angle coverage |
extract | Fetch 1-20 URLs and return clean content, labels, and optional highlights | summarization, RAG, fact lookup |
image_search | In Beta — contact us for beta access. Search the web for images by text query (optional reference image_url) | finding pictures, photos, visual references |
video_search | In Beta — contact us for beta access. Search the web for videos by text query | finding videos, clips, footage |
Reference docs:
- Search: docs.octen.ai/api-reference/search
- Extract: docs.octen.ai/api-reference/extract
Keep the tools always on (optional)
In clients with MCP tool search enabled (the Claude Code default), tools are
deferred — the model runs a ToolSearch step to load them on demand. If you'd
rather have the Octen tools resident from the first turn (no discovery step), set
alwaysLoad on the server in your .mcp.json (Claude Code v2.1.121+):
{
"mcpServers": {
"octen": {
"command": "npx",
"args": ["-y", "octen-mcp"],
"env": { "OCTEN_API_KEY": "your-key-here" },
"alwaysLoad": true
}
}
}
Each always-loaded tool uses context on every turn, and alwaysLoad blocks startup
until the server connects (capped at the ~5s connect timeout), so reserve it for tools
you hit constantly. To keep the cost down, mark just the highest-traffic tools — e.g.
search and broad_search — with "anthropic/alwaysLoad": true in each tool's _meta,
leaving the rest deferred.
Why agents like this
Most extract tools stop at "here is the page body." Octen helps one step earlier:
- Skip bad pages early:
page_structure.primary == "No Main Content"tells the agent it hit a login wall, empty shell, or similar non-content page. - Filter by topic early:
categoryhelps a pipeline ignore pages outside the target vertical before embedding or summarizing. - Use less context:
queryreturnshighlightswhen the user wants a specific fact instead of the full page.
For the full decision tree and integration patterns, see docs/best-practices.md.
Example prompts
Fetch octen.ai and summarize the main product features.Search for recent MCP news from the last week.Fetch these URLs and only summarize the ones whose category is Finance.Search site:docs.anthropic.com prompt caching and return only the relevant highlights.
Environment variables
| Variable | Required | Default | Notes |
|---|---|---|---|
OCTEN_API_KEY | yes | — | |
OCTEN_API_URL | no | https://api.octen.ai | |
OCTEN_ENABLE_BETA_TOOLS | no | on | Set to false/0/off/no to hide the Beta image_search / video_search tools from discovery. |
HTTPS_PROXY / HTTP_PROXY / NO_PROXY | no | — | Honoured since 0.4.0. Node's built-in fetch ignores these by default, so before 0.4.0 the server could not reach the API from behind a proxy even when every other tool on the machine could. Set them explicitly in your client config — see below; most MCP clients do not pass your shell environment through. |
OCTEN_KEEP_ALIVE_MS | no | 60000 | How long an idle connection is kept for reuse. undici's own default of 4s meant nearly every call re-paid a full TLS handshake (~515ms measured). 60s is measured against api.octen.ai, which closes idle connections between 60s and 90s — staying under that means we always release first, instead of dispatching onto a socket the origin has already closed. Re-measure if you point OCTEN_API_URL elsewhere. |
OCTEN_KEEP_ALIVE_MAX_MS | no | 600000 | Upper bound on the above when the origin advertises its own Keep-Alive hint. api.octen.ai does not send one. |
OCTEN_CONNECT_TIMEOUT_MS | no | 10000 | Ceiling on establishing the outbound connection to api.octen.ai — unrelated to the MCP client's own startup connect timeout mentioned above. Lower it (e.g. 5000) on a path where connections fail intermittently, so the automatic retry engages sooner. |
OCTEN_RETRY | no | on | Set to false/0/off/no to disable the single automatic retry on connection-level failures. Retries cost quota when the original request had in fact reached the server. |
OCTEN_HTTP2 | no | off | Opt into HTTP/2. Measured no faster for the usual one-request-at-a-time pattern, and not reliable through every CONNECT proxy — worth trying if you issue many tool calls in parallel. |
OCTEN_MCP_DEBUG | no | off | Request tracing on stderr (stdout carries MCP framing). See below. |
Setting these in Claude Desktop (and where the logs go)
MCP servers do not inherit your shell environment. Claude Desktop spawns them
with HOME, LOGNAME, PATH, SHELL and USER — and nothing else except what
you put in the server's env block. A proxy configured system-wide will not be
picked up; it has to be named explicitly, alongside the API key:
{
"mcpServers": {
"octen": {
"command": "npx",
"args": ["-y", "octen-mcp"],
"env": {
"OCTEN_API_KEY": "your-key-here",
"HTTPS_PROXY": "http://proxy.example:8080"
}
}
}
}
Restart Claude Desktop after editing the config — it reads it at launch.
Server output lands in:
- macOS:
~/Library/Logs/Claude/mcp-server-octen.log - Windows:
%APPDATA%\Claude\logs\mcp-server-octen.log
Everything the server writes to stderr, including the tracing below, goes there.
Diagnosing a slow or failing call
Add "OCTEN_MCP_DEBUG": "1" to the env block above while you are investigating,
and take it out afterwards — the client appends this to a log file that is
never rotated, so leaving it on grows that file for the life of the install.
With it on, every call is traced to stderr:
[octen-mcp 2026-08-14T04:36:36.219Z] call #1 received tool=search
[octen-mcp 2026-08-14T04:36:36.637Z] connect #1 established to api.octen.ai in 410ms peer=203.0.113.10:443 tls=TLSv1.3 alpn=http/1.1
[octen-mcp 2026-08-14T04:36:37.155Z] /search attempt=1 status=200 elapsed=935ms socket=new request_id=42cd56a5-…
[octen-mcp 2026-08-14T04:36:37.157Z] call #1 returning tool=search handler_total=938ms
Each field answers a specific question:
call #N receivedtimestamp — when the call reached this process. Subtract it from the time your MCP client issued the tool call: the difference is time spent entirely outsideocten-mcp, in the host or in whatever relays between them. A client-side stopwatch alone cannot separate that from time we are responsible for.connect … established in Xms— a handshake happened, and what it cost.connect FAILEDnames the phase and error code instead.peer=/tls=/alpn=— which address the connection actually reached (the API hostname is anycast, so the hostname alone cannot tell you which edge), and the negotiated TLS version and protocol — a mismatch there otherwise presents as an unexplained slow or failed handshake.socket=new/socket=reused— whether this call paid for a handshake. This is the difference between "the service is slow" and "the connection was thrown away between calls".elapsedvshandler_total— time in the HTTP request vs time in the tool handler. A large gap means the cost is in request assembly or response formatting, not the network.request_id— in this trace only: the client-generated correlation id, stable across the retry, tying a call's attempts together. It never appears in user-facing error messages, deliberately: Octen support cannot look it up (the gateway does not record the header), and an id labelledrequest_idreads like one they could. Error messages carry only ids verified searchable on Octen's side — today that is exactly one: the server's ownrequest_idfrom an API error envelope.
Failures name the cause rather than fetch failed:
Network error calling Octen Search: code=ECONNREFUSED cause=connect ECONNREFUSED 203.0.113.9:443
address=203.0.113.9:443 — could not establish a connection.
If this machine requires an HTTP proxy, set HTTPS_PROXY.
UND_ERR_CONNECT_TIMEOUT means the connection was never established, ECONNRESET
means it was established and then torn down, and ENOTFOUND means DNS — three
different problems with three different owners.
Request timeouts: search and the media tools default to 30s, broad_search to 120s (raisable to 300s),
and extract to its per-URL budget plus headroom. The search tools accept a
timeout parameter to override; extract's timeout is the server-side,
per-URL fetch budget, so the client ceiling is derived from it rather than
equal to it. The automatic retry draws down the same deadline as the first
attempt, so the stated timeout bounds the whole call, retry included.
Local development
git clone https://github.com/Octen-Team/octen-mcp.git
cd octen-mcp
npm install
npm run build
OCTEN_API_KEY=<key> npm run inspect
More docs
- Best practices for agent integration: docs/best-practices.md
- Search API reference: docs.octen.ai/api-reference/search
- Extract API reference: docs.octen.ai/api-reference/extract
License
MIT © Octen
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
MarkItDown
Freeby Microsoft · Content & Media
Convert files (PDF, Word, Excel, images, audio) to Markdown for LLM consumption
MCP Marketplace
Freeby mcp-marketplace · Developer Tools
Search and install MCP servers from inside your AI client.
FinAgent
Freeby mcp-marketplace · Finance
Free stock data and market news for any MCP-compatible AI assistant.
