Geonode logo

Unlimited Google Scholar scraper for AI agents.

A Google Scholar scraper for search results, cited-by lists, author profiles, versions and case law. Unlimited requests for one flat bill from $50/mo, no credits, and a blocked request is free.

Start 48-hour free trial

No card. Unlimited pages. 5 threads. Failed requests $0

Unlimited web scraper
01

Images

02

Title

  • Title
  • Condition
03

Price & bidding

  • Price
  • Bids
  • Time left
  • Watchers
04

Seller

  • Seller rating
  • Sold listings
05

Specifics & shipping

  • Item specifics
  • Shipping
  • Item location

JavaScript rendered, parsed to JSON or Markdown, and a fetch that fails after three retries is never billed

Call it on Google Scholar

curl -X POST "https://scraper.geonode.io/v1/extract" \
  -H "X-Api-Key: $GEONODE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "url": "https://scholar.google.com/scholar?q=attention+is+all+you+need&hl=en",
  "formats": [
    "json",
    "markdown"
  ],
  "render_js": true,
  "proxy": {
    "country": "US"
  }
}'

You came for an Google Scholar scraper. The key opens everything.

100% first-pass success

you’re here

Result entry: Title, Authors, Venue, Year, Snippet, Result ID, Document type tag

Result links: Cited by count, Cited by URL, Related articles URL, All versions count, PDF link, Publisher link

Search controls: Query, Year range, Sort by date flag, Include citations flag, Language, Start offset

Author profile: Author name, Affiliation, Verified email domain, Interests, Total citations, h-index, i10-index

Profile publications: Publication title, Co-authors, Journal, Year, Citations per paper, Citations per year chart

One subscription, from $50 a month flat. Everything after the first line adds $0

The hard part is getting in. That's our part.

  • $50/moflat, priced by concurrency, not per request

  • 0credits, meters or per-page charges

  • $0for any Google Scholar request that still fails after 3 retries

  • 48 hfree trial on your own Google Scholar URLs

  • Cloudflare98%

  • Akamai97%

  • Imperva Incapsula97%

  • DataDome96%

  • PerimeterX / HUMAN95%

  • reCAPTCHA v395%

  • Kasada94%

  • hCaptcha93%

First-pass success by anti-bot vendor, measured by us over 30 days, so of course it flatters us. The 48-hour trial exists so you can count against your own URLs. Full benchmark

What unlimited changes for Google Scholar data

  • Sold-comps pricing

    Completed listings are the real price of anything. Pull them at scale and price with evidence.

  • Arbitrage sourcing

    Sweep categories for mispriced inventory continuously instead of when the credit budget allows

  • Demand research

    Bids, watchers and sell-through, across a whole category, on one flat bill

  • Collection tracking

    Every listing of the thing you care about, the moment it appears

Cheap is supposed to be the trade-off

  • Michael T.

    Switched from Bright Data. Same quality, fraction of the cost. Their wholesale model actually makes sense.

  • Kristina Halvorson

    Finally a proxy provider that actually feels built for developers

  • Edgar Weissnat

    We run multiple large-scale data collection pipelines across several regions, and proxy reliability has always been one of the biggest bottlenecks. After switching to Geonode we noticed two immediate improvements: connection stability and predictable pricing. Previously we had to constantly optimize traffic to avoid massive proxy bills. With Geonode’s wholesale pricing model we can focus on building products instead of worrying about proxy usage. Integration with our Python stack and Playwright automation was straightforward and took less than a day.

  • Gladys Paucek

    Best proxy infrastructure I've used. The MCP integration with our AI agents was seamless.

  • Gladys Paucek

    Finally something that doesn’t feel like enterprise sales software.

Pick your thread count. That's the only decision.

Nothing is metered. Plans price threads (how many requests run at once), and pages stay unlimited on every one. An Google Scholar page that needs an anti-bot bypass costs exactly what a plain page costs

Roughly what each plan clears in a month, by endpoint

Starter

2 threads

$50 /mo

~470K
~280K
~1.1M
~610K
~120K
~400K

Growth

5 threads

$100 /mo

Popular
~1.2M
~700K
~2.8M
~1.5M
~300K
~1M

Scale

25 threads

$400 /mo

~5.9M
~3.5M
~13.8M
~7.6M
~1.5M
~5M

Pro

100 threads

$1,250 /mo

~23.5M
~14M
~55M
~30M
~5.9M
~20M

Max

250 threads

$2,500 /mo

~59M
~35M
~137M
~76M
~15M
~50M

The 48-hour trial runs at five threads, no card. Custom terms past 250 threads.

Yeah, but...

We do not yet publish a benchmark number for scholar.google.com, so you will not find a pass rate here. The defence is the one every Google Scholar scraper hits: after a small burst of queries from one address Scholar returns a 429 with a reCAPTCHA page titled "We're sorry", and the block sticks to that IP and cookie jar for hours. The Google Scholar scraper rotates residential IPs from our own pool, keeps query rate per address low, and treats the captcha page as a failure rather than a result. A request that fails after three retries is never billed, on any plan. Test it on your own queries in the 48-hour trial.

Search results at /scholar?q= with year range, sort by date and include-citations controls; cited-by lists at /scholar?cites={id}; related articles; the all-versions list for one paper; author profiles at /citations?user={id} with h-index, i10-index, the citations-per-year chart and the paginated publication table; and case law results when as_sdt is set. Each has its own target_param so the Google Scholar scraper applies the right parser and returns a stable JSON object. Every result carries the Scholar result ID, the publisher link and the direct PDF link when Scholar shows one, so the Google Scholar scraper output joins cleanly to a citation graph.

The cite popup on each result is a separate request to /scholar?q=info:{id}:scholar.google.com/&output=cite, and the BibTeX link behind it is another; both work for a logged-out visitor, and the Google Scholar scraper returns the BibTeX, MLA, APA and Chicago strings for a result ID in one object. Cited-by is a search like any other, so you pass the cites parameter and page through it with start, then recurse one level at a time to build a graph. Scholar renders results server-side, so rendering is off by default for this target, which keeps the Google Scholar scraper fast; turn render_js on only for the citations-per-year chart on profiles.

Scholar results, cited-by lists and public author profiles are visible to a logged-out visitor, and in the US the hiQ v. LinkedIn decisions treat collecting public pages as outside the Computer Fraud and Abuse Act; Google's Terms of Service still forbid automated queries, so contract risk is separate from criminal risk. The Google Scholar scraper never signs in, so My Library, alerts and private profiles are out of scope. Scholar caps any query at about 1,000 results, the start offset stops at 990, and each page holds at most 20 entries; split a query by year to go deeper. Author names and affiliations are personal data under GDPR, so store what you need.

No request credits and no per-page meter. Plans are priced by concurrency, which is the number of requests in flight at once, from $50 a month. Run as many requests through that concurrency as you can, around the clock, for the same bill. A request that fails after three retries is not counted against anything.

Yes. Pass a country code and the request routes through a residential IP there, so marketplaces return local prices and currency, search engines return local results, and geo-restricted pages open. City and ASN targeting are available on the residential pool where a site is sensitive to it.

Typed JSON with the field table on this page, clean Markdown for feeding an LLM context window, or raw rendered HTML if you run your own parser. For pages without a dedicated parser you can pass a JSON schema and get AI-extracted fields back in that shape.

48 hours. Your worst Google Scholar URLs.

Five threads, no card, and the clock starts at your first API call. If it doesn’t earn the switch, don’t switch.