Server data from the Official MCP Registry
Web scraping for agents: scrape, crawl, map, search and extract pages as clean markdown or JSON.
About
Web scraping for agents: scrape, crawl, map, search and extract pages as clean markdown or JSON.
Remote endpoints: streamable-http: https://api.snoopscan.com/mcp-oauth streamable-http: https://api.snoopscan.com/mcp
Security Report
Valid MCP server (1 strong, 2 medium validity signals). No known CVEs in dependencies. Imported from the Official MCP Registry.
Endpoint verified · Requires authentication · 1 issue found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Found in Source Code
Found by scanning the linked source code. This listing connects to a hosted endpoint, so none of this runs on your machine: it describes what the server software does where it is hosted.
How to Connect
Remote Plugin
No local installation needed. Your AI client connects to the remote endpoint directly.
Add this to your MCP configuration to connect:
{
"mcpServers": {
"com-snoopscan-snoopscan": {
"url": "https://api.snoopscan.com/mcp-oauth"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
SnoopScan
Turn any public page into clean markdown, structured JSON, or a schema you define — including the pages that block everything else. An MCP server for AI agents, a REST API, and SDKs.
This repository is SnoopScan's open core. It runs on its own, and the hosted API at snoopscan.com adds the parts that are not here. See Open source vs hosted API.
Set up in your AI chat
Paste this into Claude Code, Codex, Cursor, VS Code or any AI agent:
Read and follow https://snoopscan.com/agent-onboarding/SKILL.md
Your agent reads the guide and does every step. When your browser opens, sign in or create a free account (1,500 credits a month, no card) and press Approve. That's it: ask it to read a page.
In Claude Code you can install the SnoopScan plugin instead, which brings SnoopScan's skills and its MCP server. Type these in Claude Code's chat:
/plugin marketplace add SnoopScan/snoopscan-core
/plugin install snoopscan@snoopscan
The Claude app and ChatGPT connect by signing in, with no key to copy: add a
custom connector (Claude) or a developer-mode app (ChatGPT) with the address
https://api.snoopscan.com/mcp-oauth. Steps for every client are at
snoopscan.com/integrations.
Use it from your code
The examples below call the hosted API and read your key from
SNOOPSCAN_API_KEY. Create one at
snoopscan.com/app/keys, or let your agent
fetch it for you with the setup above.
No key yet? /v1/scrape, /v1/search and /v1/parse also work with no
Authorization header: 20 free calls a day per IP address, for pages that load
without a browser. Every answer says how many free calls are left.
curl
curl -X POST https://api.snoopscan.com/v1/scrape \
-H "Authorization: Bearer $SNOOPSCAN_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/", "formats": ["markdown", "links"]}'
Python
pip install snoopscan
import os
from snoopscan import SnoopScan
snoop = SnoopScan(api_key=os.environ["SNOOPSCAN_API_KEY"])
print(snoop.scrape("https://example.com/").markdown)
JavaScript / TypeScript
npm install snoopscan
import { SnoopScan } from 'snoopscan';
const snoop = new SnoopScan({ apiKey: process.env.SNOOPSCAN_API_KEY! });
console.log((await snoop.scrape('https://example.com/')).markdown);
Other languages
Every SDK has the same methods, the same error codes and the same retry rules,
and reads SNOOPSCAN_API_KEY from the environment when no key is passed. Each
directory's README covers every method.
| Language | Install | Source |
|---|---|---|
| PHP 8.1+ | composer require snoopscan/snoopscan | sdk/php |
| Go 1.22+ | go get github.com/SnoopScan/snoopscan-core/sdk/go | sdk/go |
| Ruby 3.0+ | gem install snoopscan | sdk/ruby |
| Java 11+ | com.snoopscan:snoopscan (Maven, Gradle) | sdk/java |
| .NET 8+ | dotnet add package SnoopScan | sdk/dotnet |
| Rust | cargo add snoopscan | sdk/rust |
| Elixir | {:snoopscan, "~> 0.1"} | sdk/elixir |
Features
- Clean output: markdown, HTML, links, screenshots, or JSON in your own schema.
- Handles the hard pages: JavaScript rendering, blocked pages and proxies, with no setup (hosted API).
- Whole sites: crawl a site, map its URLs, or batch thousands of URLs as one job.
- Structured data: fill a JSON schema or a named template, with a confidence score per page.
- Stores and blogs in one call: a store's whole product catalog or a blog's posts, without a crawl.
- Pay for results: failed requests cost nothing, and every response
carries a
costobject. - Agent ready: every endpoint is also an MCP tool.
Open source vs hosted API
This repository contains:
- the REST API
- extraction
- block detection
- crawling
- the MCP server
- HTTP fetching
- storage
- the client SDKs: Python, JS/TS, PHP, Go, Ruby, Java, .NET, Rust and Elixir
The hosted API at snoopscan.com adds:
- browser rendering
- proxies and anti-bot handling
- lead generation
A self-hosted copy reads pages over HTTP. A page that needs a browser comes
back as BLOCKED, never as an empty success. The hosted API's free plan
includes 1,500 credits a month.
MCP server
Every endpoint is also an MCP tool, so an agent can call them directly with no
glue code. The hosted server is https://api.snoopscan.com/mcp (streamable
HTTP, Authorization: Bearer <key>), and https://api.snoopscan.com/mcp-oauth
signs in instead of taking a key. Connected to /mcp with no key at all, an
agent can try scrape, search and parse on the same free daily trial as
REST; every other tool asks for a free key. The easiest setup is the one line above: your
agent follows the agent onboarding guide, which
covers Claude Code, Codex, Cursor, VS Code, Windsurf, the Claude app, ChatGPT
and any other client that speaks streamable HTTP.
Tools
- Pages:
scrape,fetchMore,search,parse,extract,checkChanges - Sites:
map,crawl,crawlStatus,crawlPages,listProducts,listPosts - Domains:
domain,domainReport,domainSnapshots,domainScreen - Companies and leads:
company,findContacts,people,hiring,findLeads,leadsStatus
The product and post listings, the domain report, its snapshots, the domain screen, and the company, contact, people, hiring and lead tools need the hosted API. This repository answers them with a plain "not available on this deployment".
Every tool has a title and read-only / destructive labels. None of them submits, posts or buys anything.
Limits
The tools are designed to keep an agent's context small:
- Every content tool takes a
maxCharsbudget. - Truncation is always visible and returns a continuation token.
- Crawl tools never inline page bodies.
- Guardrails cap pages per session, concurrent crawls, crawl size and bandwidth.
- Every refusal explains the limit, so an agent can adapt instead of retrying blindly.
executeJavascriptis not exposed over MCP at any level.
API endpoints
POST /v1/scrape one URL -> markdown, html, links, screenshot, json
(+ `actions`: click, type, scroll, then read the result)
POST /v1/crawl a whole site, page by page
POST /v1/map every URL on a site, ordered by relevance
POST /v1/search the web, with the results scraped
POST /v1/extract your JSON schema — or a named template — filled from the page
GET /v1/templates the named templates (product, article, jobPosting, …) and their fields
POST /v1/batch/scrape many URLs, one job
POST /v1/parse a PDF, DOCX or XLSX as markdown
POST /v1/monitor watch URLs on a schedule, webhook on change
POST /v1/products a store's whole product catalog
POST /v1/posts every post from a blog or forum
POST /v1/company enrich one company from its own site
POST /v1/domain a domain's registration, DNS and backlinks
POST /v1/domain/report a buyer's report on a domain (hosted API)
POST /v1/domain/screen which of up to 500 names can be registered (hosted API)
POST /v1/places/search local business listings, with contacts
POST /v1/leads businesses by trade and town, with their contacts (hosted API)
Every page includes metadata.platform, the platform that built it, and
every response carries a cost object.
Extract data with a schema
curl -X POST https://api.snoopscan.com/v1/extract \
-H "Authorization: Bearer $SNOOPSCAN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"urls": ["https://example.com/product/1"],
"schema": {"type": "object",
"properties": {"name": {"type": "string"},
"price": {"type": "string"}}}
}'
Monitor a page for changes
This checks the page every 60 minutes and sends a webhook when it changes.
curl -X POST https://api.snoopscan.com/v1/monitor \
-H "Authorization: Bearer $SNOOPSCAN_API_KEY" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://example.com/pricing"],
"intervalMinutes": 60,
"webhook": "https://your.app/hook"}'
Development
Set up and run the engine locally:
uv venv --python 3.12 && uv pip install -e ".[dev]"
uv pip install -e sdk/python --python .venv/bin/python
cp .env.example .env && createdb scraping_engine && .venv/bin/alembic upgrade head
.venv/bin/python tools/create_key.py "local-dev" --rpm 600
.venv/bin/uvicorn engine.api.app:app --reload --port 8099
The JS/TS SDK is in sdk/js and builds on its own:
cd sdk/js && npm install && npm run build && npm test
The other SDKs build the same way, each in its own directory with its own
toolchain, and each README's Development section has the command. Their
dependency licenses are checked by sdk/check_licences.py:
python3 sdk/check_licences.py rust # after the SDK's own tests have run
Run the tests, lint and type checks:
.venv/bin/pytest engine/tests -q
.venv/bin/ruff check engine tools && .venv/bin/mypy --strict engine/core
Dependency licenses
Every dependency must be MIT, Apache-2.0, BSD, ISC or MPL-2.0. The check is blocking: a build that pulls in any other license fails. It also runs weekly, because a transitive dependency can change its license.
.venv/bin/pip-licenses --format=json > licences.json \
&& .venv/bin/python tools/check_licences.py licences.json
Two results of this rule are already in the code. psycopg2 (LGPL) is not
used, and Alembic migrates through asyncpg instead. tld, which comes in as a
transitive dependency, has a recorded allowlist entry that names which branch
of its tri-license we use.
Contributing
Contributions must follow three rules:
- Clean-room implementation. Do not use code from any AGPL, GPL or SSPL project. The API mirrors the option names common in this category, since interface compatibility is legitimate, but the implementation is independent.
- The license check is blocking, not advisory.
- No personal or identifying data in the repo. Configuration comes from
environment variables only. Every fixture uses
example.comor.invalid.
License
The engine is licensed under AGPL-3.0. If you run a modified version as a
network service, section 13 requires you to offer its source to your users.
The running instance does this itself at GET /v1/source, so it does not
depend on a document that a fork might forget to update.
Every SDK in sdk/ is MIT. An AGPL client library would push the
copyleft into every application that imports it, which is not what a client
library is for.
Some modules are proprietary and are not covered by the AGPL: browser
rendering, proxy and anti-bot handling, and lead generation. They are listed in LICENSE-PROPRIETARY, which is generated from
engine/licensing.py. CI enforces the split: tools/check_split.py fails the
build if any public module depends on one of them.
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Git
Freeby Modelcontextprotocol · Developer Tools
Read, search, and manipulate Git repositories programmatically
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Worldmonitor
Freeby Koala73 · Developer Tools
Live markets, conflicts, country risk, chokepoints, energy, and China decision signals. 89 tools.
Paperclip
Freeby Paperclipai · Developer Tools
Trending hip-hop artist momentum scores across four cultural dimensions.
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
