Skip to content

teslashibe/web-scrape

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

3 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

web-scrape

CLI-first website scrape + anonymous browser fetch.

  • SSRF-safe public http(s) only
  • robots.txt honored (fail-open on missing/unreachable robots)
  • Protection taxonomy — blocked / captcha / Cloudflare / login / rate-limit statuses instead of garbage text
  • Optional CapSolver — Turnstile + Cloudflare “Just a moment…” clears (sticky residential proxy; CapSolver needs socks5 for Webshare)
  • Thin MCP adapter — same library over JSON-RPC HTTP

Does not bypass protections by default. CapSolver is opt-in via env. No social-login profiles or cookies are reused.

Install

git clone https://github.com/teslashibe/web-scrape.git
cd web-scrape
npm install
npx playwright install chromium   # for `fetch` / browser path

Node ≥ 20. Uses undici v7 (compatible with Node 20).

CLI (preferred for agents)

# Readable browser fetch (Playwright)
./bin/web-scrape.mjs fetch https://example.com/ --json

# Structured single-page scrape (HTTP + HTML extract)
./bin/web-scrape.mjs scrape https://example.com/ --json

# MCP HTTP server (default :8091)
./bin/web-scrape.mjs serve --port 8091

Exit codes: 0 ok · 2 structured non-ok status · 1 usage/error.

Library

import { scrape, browserFetchURL } from "web-scrape";

const page = await browserFetchURL({ url: "https://example.com/" });
const brief = await scrape({ url: "https://example.com/" });

Docker / Kubernetes

docker build -t web-scrape .
docker run --rm -p 8091:8091 \
  -e WEB_FETCH_TURNSTILE_PROVIDER=capsolver \
  -e WEB_FETCH_TURNSTILE_API_KEY=CAP-... \
  -e WEBSHARE_RESIDENTIAL_USER=... \
  -e WEBSHARE_RESIDENTIAL_PASS=... \
  -e WEBSHARE_USERNAME_TEMPLATE='{user}-{country}-1' \
  web-scrape

Expose Service port 8091. Probe GET /mcp/v1/ready (200 ready / 503 not_ready).

MCP

POST /mcp/v1          JSON-RPC 2.0 (initialize, tools/list, tools/call)
GET  /mcp/v1/health   liveness + limits + metrics
GET  /mcp/v1/ready    readiness

Tools:

Tool Purpose
browser_fetch_url Ephemeral Playwright fetch → readable text or structured status
scrape_website_context Single-page HTTP scrape → title/summary/product/audience fields

CapSolver + proxy (optional)

Env Purpose
WEB_FETCH_TURNSTILE_PROVIDER=capsolver Enable provider
WEB_FETCH_TURNSTILE_API_KEY CapSolver key
WEBSHARE_RESIDENTIAL_USER / PASS Residential proxy
WEBSHARE_USERNAME_TEMPLATE Prefer sticky {user}-{country}-1 (required for cf_clearance IP affinity)
WEB_SCRAPER_PROXY_URL Full proxy URL override
BROWSER_FETCH_TIMEOUT_MS Playwright deadline (default rises when CapSolver+proxy set; max 180s)

CapSolver AntiCloudflareTask is sent as socks5:host:port:user:pass (no page html — CapSolver rejects it as invalid html). Playwright egress stays HTTP proxy. Managed “Just a moment…” pages use AntiCloudflareTask even when a Turnstile iframe is present. Never log proxy credentials or API keys.

Tests

npm test          # unit fixtures (no live network)
npm run validate  # loopback MCP discovery/call
npm run smoke     # status-matrix smoke

License

Apache-2.0. CapSolver / Webshare are optional third-party paid services — you bring your own keys.

About

CLI-first website scrape + browser fetch with SSRF/robots guards and optional CapSolver Cloudflare clears. MCP adapter included.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages