# Geonode Documentation — Single File (llms-full.txt) Generated: 2026-08-17 (v2, includes full endpoint schemas). This file concatenates the Geonode Scraper API documentation for LLM and AI-agent consumption. Canonical docs: https://docs.geonode.com. Machine-readable site index: https://geonode.com/llms.txt Agent onboarding skill: https://geonode.com/agent-onboarding/SKILL.md --- # API Overview Source: https://docs.geonode.com/docs/scraper-api The Geonode Scraper API helps you extract, discover, and process web content without managing browsers, proxies, or scraping infrastructure. Send a URL and receive structured content as Markdown or HTML. The API also supports JavaScript rendering, geo-targeting, batch processing, website crawling, and webhook notifications. ## What You Can Do With the Scraper API, you can: - Extract content from webpages - Process multiple URLs in batch jobs - Crawl websites and discover pages - Receive webhook notifications when jobs complete - Use geo-targeted proxy routing - Extract content from JavaScript-powered websites - Retrieve links found on webpages ## Available APIs | API | Use When | | ---------- | -------------------------------------------------------- | | Extraction | You already know the URL and want to extract its content | | Batch | You have multiple URLs that need to be processed | | Crawl | You want to discover and extract pages across a website | | Webhooks | You want to receive notifications when jobs complete | ## When to Use the Scraper API Use the Scraper API when you need: clean Markdown or HTML output, JavaScript rendering, geo-targeted extraction, managed proxy infrastructure, batch processing, website crawling, asynchronous processing, webhook notifications. If you need complete control over browser automation, request handling, or custom scraping logic, consider using the Geonode Proxy API instead. --- # Quick Start Guide Source: https://docs.geonode.com/docs/scraper-api/quick-start This guide helps you get started with the Geonode Scraper API in a few minutes. ## Get Your API Key 1. Sign in to your Geonode account. 2. Open the Dashboard. 3. Navigate to the API Keys section. 4. Create or copy an existing API key. ## Authentication All Scraper API requests require the `X-Api-Key` header. curl -H "X-Api-Key: YOUR_API_KEY" Replace `YOUR_API_KEY` with your actual API key. ## Choose an API | API | Use When | | ---------- | ------------------------------------------------------- | | Extraction | You want to extract content from one or more webpages | | Batch | You have multiple URLs to process in a single job | | Crawl | You want to discover and extract pages across a website | | Webhooks | You want notifications when jobs complete | --- # Scraper API Reference Source: https://docs.geonode.com/docs/scraper-api/reference Base URL, authentication, output formats, and endpoint overview for the Geonode Scraper API. ## Base URL https://scraper.geonode.io ## Authentication Send your Scraper API key in the `X-Api-Key` header. X-Api-Key: YOUR_API_KEY Keep your API key private. Do not expose it in frontend code, public repositories, logs, screenshots, or support messages. ## Endpoints | Method | Endpoint | Description | | -------- | ------------------------------------------- | ----------------------------------------------------------------------- | | GET | /health | Check whether the Scraper API service is healthy. | | POST | /v1/extract | Extract Markdown and/or HTML from a single webpage. | | GET | /v1/extract/{job_id} | Poll one async extraction job and retrieve the result when it is ready. | | GET | /v1/extract/jobs | List and filter previous extraction jobs. | | POST | /v1/batch | Start an asynchronous extraction job for multiple URLs. | | GET | /v1/batch/{job_id} | Poll a batch job and retrieve paginated item results. | | DELETE | /v1/batch/{job_id} | Cancel a batch job that is still running. | | POST | /v1/map | Discover URLs from a base URL using sitemap and HTML link discovery. | | GET | /v1/statistics | Retrieve aggregated extraction statistics for a date range. (Coming soon) | | POST | /v1/crawl | Start a site crawl from one seed URL. | | GET | /v1/crawl/{job_id} | Poll a crawl job and retrieve paginated page results. | | DELETE | /v1/crawl/{job_id} | Cancel a crawl job that is still running. | | POST | /v1/webhooks | Register a webhook subscription. (Coming soon) | | GET | /v1/webhooks | List registered webhooks. | | GET | /v1/webhooks/{webhook_id} | Retrieve one webhook subscription. | | PATCH | /v1/webhooks/{webhook_id} | Update a webhook subscription. | | DELETE | /v1/webhooks/{webhook_id} | Delete a webhook subscription. | | POST | /v1/webhooks/{webhook_id}/rotate-secret | Rotate a webhook signing secret. | | GET | /v1/webhooks/{webhook_id}/deliveries | List delivery attempts for a webhook. | ## Output Formats The API response envelope is always JSON. The extracted page content can be returned as Markdown, HTML, or both. {"formats": ["markdown"]} {"formats": ["html"]} {"formats": ["markdown", "html"]} Use Markdown when you need readable page content for LLM workflows, indexing, review, or text processing. Use HTML when you need structure closer to the original page. There is no structured JSON content extraction format in the current public POST /v1/extract schema. ## Common Options | Option | Use it when | | ----------------- | ------------------------------------------------------------------------------------ | | render_js | The returned content looks like a shell, loading state, or navigation-only document. | | proxy | You need a specific country or proxy type. | | processing_mode | You want to use async mode for slow pages. | | headers | You need to send additional target-request headers. | | extract_links | You want links found on the extracted page included in the response. | --- # Endpoint Schemas Source: https://docs.geonode.com/docs/scraper-api/v1/* (per-endpoint pages) Production status: Extraction, Batch, Crawl, Map, and Health are live. Webhooks and Statistics are documented as "Coming soon — not yet available in production"; their contracts below reflect planned behavior. ## POST /v1/extract — Extract Content Source: https://docs.geonode.com/docs/scraper-api/v1/extraction/extract-content Extract clean Markdown or HTML from any URL with sync or async mode. Auth: X-Api-Key header. Request body (application/json): - url (required, uri) — URL to extract content from - formats (array<"markdown"|"html">, default ["html"]) — output formats - render_js (boolean, default false) — use a headless browser to render the page (slower, more expensive) - processing_mode ("sync"|"async", default "sync") — sync blocks until done; async returns a job ID - proxy (object|null, default {"type":"residential"}) — when omitted or null, defaults apply: residential type, country auto-resolved from the target URL (domain/path overrides first, then the TLD; if nothing matches, the upstream proxy provider picks the exit region). Fields: country (ISO-2), type ("residential"|"datacenter"|"mix") - headers (object|null) — custom HTTP headers for the target request - wait_config (object|null) — grouped browser wait policy: wait_until ("commit"|...), wait_for (selector), wait_timeout (ms). When omitted/null, browser requests use the adaptive default settle policy. A non-null wait_config opts into caller-owned explicit wait mode and, if render_js is not set, auto-enables browser rendering. Sending render_js=false together with non-null wait_config is rejected as ambiguous. Responses: - 200 (sync success): { "data": {"html", "markdown"}, "metadata": {"url", "render_js", "http_status", "duration_ms", "formats", "proxy": {"country","type"}, "processing_mode", "headers", "wait_config"}, "tokens_charged": N } - 202 (async accepted): { "job_id", "status": "queued", "status_url", "estimated_tokens" } - 4xx/5xx: { "code", "message", "correlation_id", "retryable" } or extraction error { "error": {"code","message","retryable","details"}, "tokens_charged" } `tokens_charged` = number of requests charged (schema field name is historical). ## GET /v1/extract/{job_id} — Get Extraction Job Source: https://docs.geonode.com/docs/scraper-api/v1/extraction/get-extraction-job Poll one async extraction job by job_id (uuid). Returns the same response envelope as a sync extract once the job completes; job status while pending. ## GET /v1/extract/jobs — List Extraction Jobs Source: https://docs.geonode.com/docs/scraper-api/v1/extraction/list-extraction-jobs List and filter previous extraction jobs for the authenticated key. ## POST /v1/batch — Start a Batch Job Source: https://docs.geonode.com/docs/scraper-api/v1/batch/start-batch-job Queue 1–1,000 URLs for asynchronous extraction in one job. Request body: - urls (required, array, 1..1000 items) - ignore_invalid_urls (boolean, default true) — invalid URLs are skipped and returned in invalid_urls instead of failing the request - formats (array, default ["markdown"]) - render_js (boolean, default false) — applies to every batch item - proxy (ProxySettings) — proxy configuration - headers (object|null) - wait_config (object|null) — applied to every batch item; same auto-enable and ambiguity rules as extract Response 202: { "job_id", "status": "queued", "status_url", "accepted_urls", "invalid_urls": [...] } ## GET /v1/batch/{job_id} — Get Batch Job Status Source: https://docs.geonode.com/docs/scraper-api/v1/batch/get-batch-job-status Poll a batch job; results are paginated. Query: page (default 1), page_size (default 10, max 50). Response 200: { "job_id", "status", "batch_config": {"render_js","formats","proxy","headers","wait_config"}, "token_summary": {"tokens_charged_total","tokens_reserved"}, "total_urls", "completed_urls", "failed_urls", "cancelled_urls", "pending_urls", "invalid_urls", "created_at", "completed_at", "results": [{ "input_index", "url", "status", "error_code", "error_message", "data": {"html","markdown"}, "metadata": {"http_status","duration_ms","tokens_charged"} }] } ## DELETE /v1/batch/{job_id} — Cancel a Batch Job Source: https://docs.geonode.com/docs/scraper-api/v1/batch/cancel-batch-job Cancel a running batch job. 202 accepted; 409 if the job cannot be cancelled in its current state. ## POST /v1/crawl — Start a Crawl Job Source: https://docs.geonode.com/docs/scraper-api/v1/crawl/start-crawl-job Crawl a website from a seed URL up to a given depth and page limit. Request body: - url (required, uri) — seed URL - depth (int, default 2, range 1..10) — maximum BFS depth from the seed (1 = seed only) - limit (int, default 50, range 1..10000) — maximum pages to crawl - formats (array, default ["markdown"]) — per page - render_js (boolean, default false) - same_domain_only (boolean, default true) — only follow links on the seed's domain - include_subdomains (boolean, default false) — with same_domain_only, also include subdomains - proxy (object|null, default {"type":"residential"}) — per-page defaults as in extract - wait_config (object|null) — applied to every crawled page; same rules as extract Response 202: { "job_id", "url", "status": "queued", "status_url", "estimated_pages" } Each crawled page counts as one request. ## GET /v1/crawl/{job_id} — Get Crawl Job Status Source: https://docs.geonode.com/docs/scraper-api/v1/crawl/get-crawl-job-status Poll a crawl job; results are paginated. Query: page (default 1), page_size (default 10, max 50). Response 200: { "job_id", "url", "status", "crawl_config": {"render_js","formats","same_domain_only","include_subdomains","proxy","wait_config"}, "token_summary": {"tokens_charged_total","tokens_reserved"}, "total_pages", "completed_pages", "failed_pages", "cancelled_pages", "created_at", "completed_at", "results": [{ "url", "parent_url", "depth", "status", "error_code", "error_message", "links": [...], "data": {"markdown","html"}, "metadata": {"http_status","duration_ms","tokens_charged"} }] } ## DELETE /v1/crawl/{job_id} — Cancel a Crawl Job Source: https://docs.geonode.com/docs/scraper-api/v1/crawl/cancel-crawl-job Cancel a running crawl job. 202 accepted; 409 if not cancellable. ## POST /v1/map — Map URLs Source: https://docs.geonode.com/docs/scraper-api/v1/map/map-urls Returns the list of URLs found under the given base URL by combining sitemap parsing with HTML link extraction from the seed page. One successful map counts as one request. Request body: - url (required, uri) — base URL to discover links from - search (string|null) — case-insensitive substring filter applied to discovered URLs. Does NOT query a search engine; only narrows sitemap/HTML discovery results - include_subdomains (boolean, default true) — probe common sibling subdomains (docs, blog, help, support) for their own sitemaps - ignore_query_parameters (boolean, default true) — strip query parameters when normalizing discovered URLs Response 200: { "success": true, "links": [{ "url", "source": "sitemap"|... }], "warning" } 408 if URL discovery times out. ## GET /health — Health Check Source: https://docs.geonode.com/docs/scraper-api/v1/system/health-check Check whether the Scraper API service is healthy. No auth-sensitive data. ## GET /v1/statistics — Get Statistics (Coming soon) Source: https://docs.geonode.com/docs/scraper-api/v1/statistics/get-statistics Documented but not yet available in production; contract reflects planned behavior. Query: start_date (default 30 days ago), end_date (default now). Planned response: { "extraction_count", "previous_period_extraction_count", "success_rate", "average_extraction_duration", "period_tokens_used", "requests": [{"date","success_request_count","failed_request_count"}], "tokens_used": [{"date","tokens_used"}] } ## /v1/webhooks — Webhooks (Coming soon) Source: https://docs.geonode.com/docs/scraper-api/v1/webhooks/create-webhook Documented but not yet available in production; contracts reflect planned behavior. Planned event types: "extract_completed", "batch_completed", "crawl_completed". Planned endpoints: POST /v1/webhooks (create; returns signing secret), GET /v1/webhooks (list), GET/PATCH/DELETE /v1/webhooks/{webhook_id}, POST /v1/webhooks/{webhook_id}/rotate-secret, GET /v1/webhooks/{webhook_id}/deliveries. For launch notification or early access: https://geonode.com/contact --- # Error Handling Source: https://docs.geonode.com/docs/scraper-api/error-handling The Scraper API returns standard HTTP status codes and JSON error bodies. Handle both validation errors and extraction errors in your client, because a request can fail before extraction starts or after the API tries to process the target page. ## HTTP Status Codes | HTTP status | Meaning | | ----------- | ----------------------------------------------------------------------------------------------- | | 200 | Synchronous extraction, map request, job lookup, statistics request, webhook lookup, or health check succeeded. | | 201 | Webhook subscription was created. | | 202 | Async extraction, batch, or crawl job was accepted, or a running job cancellation was accepted. | | 204 | Webhook subscription was deleted. | | 400 | Invalid request. | | 401 | API key is missing or invalid. | | 402 | Payment required or insufficient request balance. | | 404 | Job, webhook, or other requested resource was not found. | | 408 | Map request timed out during URL discovery. | | 409 | Batch or crawl job cannot be cancelled in its current state. | | 422 | Validation error or extraction failed. | | 429 | Request throttled or work concurrency limit reached. | | 500 | Internal server error. | | 502 | Billing service returned an upstream error. Retryable depends on the billing error. | | 503 | Service or billing service temporarily unavailable. | | 504 | Synchronous extraction timed out waiting for the browser worker. Retryable. | ## Validation Error Validation errors can return a `detail` array — the request body does not match the expected schema, such as an invalid URL or unsupported value. { "detail": [ { "type": "value_error", "loc": ["body", "url"], "msg": "Value error, URL must contain a valid hostname.", "input": "not-a-url" } ] } ## Extraction Error Extraction failures return an `error` object and `tokens_charged`. { "error": { "code": "UNPROCESSABLE_CONTENT", "message": "The target page could not be extracted.", "retryable": false, "details": null }, "tokens_charged": 0 } The response field is currently named `tokens_charged` in the API schema. In Scraper API docs and billing language, treat this value as the number of requests charged. ## Extraction Error Codes RATE_LIMITED, WORK_CONCURRENCY_LIMITED, WORK_CONCURRENCY_LEASE_EXPIRED, TEMPORARY_BLOCK, NETWORK_ERROR, CAPTCHA_CHALLENGE, PERMANENT_BLOCK, INVALID_URL, UNPROCESSABLE_CONTENT, AUTH_REQUIRED, TIMEOUT, PAYMENT_REQUIRED, INTERNAL_ERROR, PROXY_ERROR. Use the `retryable` field to decide whether a retry may help. For retryable failures, use exponential backoff and avoid retrying in a tight loop. If you receive 429, slow down the client and retry after a delay; no public numeric rate limit is specified in the current contract. --- # FAQs Source: https://docs.geonode.com/docs/scraper-api/faq Common questions about Scraper API requests, rendering, proxies, billing, and scraping behavior: What is the Scraper API? What is one credit? What happens when I run out of free requests? Do free requests roll over? Do subscription requests roll over? Do Pay as you Go requests expire? Does JavaScript rendering cost more? How is this different from using proxies directly? What sites can I scrape? Why is my result missing content? Why is the output noisy? Can I choose the country used for scraping? Which proxy types are supported? Can I pass custom headers? Can I scrape pages that require login? Should I use sync or async mode? What is the map endpoint for? Where can I see my usage? Where can I find the full API reference? Key answers in short form: JavaScript rendering does not cost extra — one request is one request. Free tier is 1,500 requests per month, renewing every month; free requests do not roll over. Pay-as-you-go requests never expire. Country targeting uses `proxy.country` (ISO-2 code); proxy types are residential, datacenter, and mix. Use sync mode for fast pages and async for slow or JavaScript-heavy pages. --- # Scraper API — Pricing Source: https://geonode.com/pricing/scraper-api · https://geonode.com/pricing.md Two pricing models — pick one: - Unlimited plans — unlimited requests, priced by concurrent threads, from $47/mo. - Per-request plans — from free (1,500 requests/month, forever) to $126/mo for 1M requests. ## Unlimited plans — priced by concurrency (monthly subscription) | Concurrent threads | Monthly requests | Price / mo | |---|---|---| | 5 (free trial) | Unlimited for 48 hours | $0 | | 2 | Unlimited | $47 | | 10 | Unlimited | $199 | | 25 | Unlimited | $399 | | 50 | Unlimited | $699 | | 100 | Unlimited | $1,799 | | Enterprise | Unlimited | Custom | ## Per-request plans — monthly subscription | Requests | Price / 1,000 requests | Price / mo | |---|---|---| | 1,500 | Free | Free, every month | | 10,000 | $0.35 | $3.50 | | 50,000 | $0.26 | $13.00 | | 250,000 | $0.17 | $43.00 | | 1,000,000 | $0.13 | $126.00 | | 3,000,000+ | Custom | Custom | ## Per-request plans — pay as you go | Requests | Price / 1,000 requests | Total | |---|---|---| | 1,500 | Free | Free, every month | | 10,000 | $0.39 | $3.90 | | 50,000 | $0.29 | $14.50 | | 250,000 | $0.19 | $47.50 | | 1,000,000+ | Custom | Custom | ## How billing works - Billed per request — no credit system, no per-request multipliers. One request = one request, regardless of JavaScript rendering or output format. - No published rate limits: no maximum requests per month, per day, or per second. - Downgrades: requests remaining on a plan you downgrade from are forfeited. - Discounts: volume tiers and annual commits (pay 10 months, get 12). ## Included JavaScript rendering; HTML, Markdown and JSON output; country targeting; batch URLs in a single call. ## Migration Switching from Firecrawl requires changing two environment variables (`SCRAPE_API_URL`, `SCRAPE_API_KEY`); the request format is compatible. --- # MCP Server Source: https://geonode.com/products/mcp One MCP server exposing web data tools — extract, batch, crawl, map, job, statistics — to AI agents. Install: `npx @geonode/mcp`. Remote: `https://scraper.geonode.io/mcp` (auth: `X-Api-Key` header). MIT licensed. Supported hosts include Claude Desktop, Claude Code, Cursor, Windsurf, Smithery and Docker MCP. The `search` tool is marked "Coming soon". The MCP Server has no separate price list. Usage draws on the same request balance as the Scraper API, at the same published rates: free tier 1,500 requests/month; unlimited plans from $47/mo by concurrency; per-request plans from $3.50/mo. One balance shared across the MCP Server, Scraper API, Map API and Crawl API. --- # Build Manifest (for the docs build step — remove from published file) This v1 was assembled manually on 2026-08-17 from the pages listed above. Endpoint schemas are now included above (v2). The automated build (GEO-3493) must additionally include, in nav order: - API Guides: Getting started, Making requests (output formats, JS rendering, waiting for dynamic content, processing modes, proxy and geo-targeting, custom headers), Jobs, Extraction, Batch, Crawl, Map, MCP (v1/scraper-mcp tree) - Code Examples; Dashboard Guides; Real World Examples (E-commerce) - Pricing pages: unlimited_scraper_api_pricing, request_based_pricing, choosing_scraper_api_plan - Proxy API docs (/docs/api-reference) and Guides (/docs/guides) - Change Log (/docs/changelog) Known issue found during assembly: FAQ answers on /docs/scraper-api/faq are not present in server-rendered HTML (accordion, questions only). Fix server-side rendering of FAQ answers — same requirement as geonode.com FAQ blocks — so both crawlers and this build can read them.