Search Provider Bench

Which web-search API leads — measured, not marketed.

Ten search APIs, 216 real queries, two AI judges that never saw which provider produced what. Quality, speed, price and reliability are scored separately — so an outage can't masquerade as bad results, and a cheap provider can't buy its way up the table.

10providers
216searches
2AI judges
snapshot date

One run, not an average. Every provider answered each of the 216 searches exactly once, from a single European host, on the snapshot date above. The spread on the charts is variation between search types — not between repeated attempts at the same query. A rerun would not land on identical numbers. Read the ordering as the finding and the exact values as one snapshot.

The table

Who wins the head-to-heads.

Every search in this benchmark was judged as a duel: two providers' results for the same query, side by side, labels stripped. A provider's win rate is simply how often it came out ahead against an average rival — 50% is the field average.

Win rate across all nine search types

Higher is better. The whisker spans each provider's best and worst category.

Linkup is absent: too little evidence to rank it in Search. It reappears — at the top — in the Extract lane below.

Speed & reliability

How fast it comes back — and whether it comes back.

Quality is half the decision. An agent that waits nine seconds on a search feels broken even when the answer is good, and a provider that drops one call in twenty forces you to write retry logic. Both charts come from the same 216 searches the judges scored.

Response time

Solid bar is the median; the pale bar behind it is the slow 5% (p95). Lower is better.

Measured from a single European host. Absolute numbers shift with your region; the ordering travels.

Searches that came back

Share of the 216 attempts that returned a usable response. Anything short of 100% is retry logic you have to write.

Win rate vs. response time

Up is better, left is faster. The corner you want is top-left.

TOP-LEFT IS BETTER

Bubble size is cost per search; the free provider is drawn as a star.

Second lane

Pulling the facts back out of a page

Finding a page and reading it are different jobs. This lane hands each provider a URL and counts how many of the target facts survive the round trip.

LOADING EXTRACT DATA

Fact recall by provider

Share of target facts recovered from the page. Higher is better.

Extract response time

Median time to fetch and return a page, 105 URLs. Pale bar is the slow 5%.

Show the full Extract table
Rank · ProviderUsable pages · Median charsFact recallCost per extract
Loading Extract results…
Cost per extract · list price where availableLoading cases…

Dropped providers: loading…

The Extract benchmark JSON could not be loaded. Serve this directory as a static site and refresh.
Price / performance

What quality actually costs.

The top of the table is not the cheapest, and the cheapest is not worthless. This is where you decide what the gap is worth to you.

Cost per search vs. average quality

Left is cheaper, up is better. The line traces the providers nothing else beats on both counts at once.

TOP-LEFT IS BETTER

Scaled to spread cheap prices; ticks show actual USD.

Provider Best-value frontier Low-cost best value Calculating best value…

Linkup isn't plotted: no rankable Search categories (insufficient evidence).

API price comparison

What each provider charges

ProviderCost per searchCost per extract
Loading prices…

List prices (vendor pay-as-you-go); real bills vary with volume.

See the row-by-row cost methodology →
The cost JSON could not be loaded. Serve this directory as a static site and refresh.
Why trust this

Three safeguards, visible in the results.

01

Blind inputs

Provider identities were hidden from the judges, so the inputs carried anonymous codes instead of provider labels.

02

Two judges

grok-4.5 and gpt-5.6-luna judged independently, so no single judge's quirks decide the outcome.

03

Confidence intervals

Every score comes with a confidence range; overlapping ranges are reported as statistical ties, not proven wins.

Limits of this release

  • Snapshot-only evidence
  • Human review of AI judges pending
  • No switching recommendation