Blind inputs
Provider identities were hidden from the judges, so the inputs carried anonymous codes instead of provider labels.
Ten search APIs, 216 real queries, two AI judges that never saw which provider produced what. Quality, speed, price and reliability are scored separately — so an outage can't masquerade as bad results, and a cheap provider can't buy its way up the table.
One run, not an average. Every provider answered each of the 216 searches exactly once, from a single European host, on the snapshot date above. The spread on the charts is variation between search types — not between repeated attempts at the same query. A rerun would not land on identical numbers. Read the ordering as the finding and the exact values as one snapshot.
Every search in this benchmark was judged as a duel: two providers' results for the same query, side by side, labels stripped. A provider's win rate is simply how often it came out ahead against an average rival — 50% is the field average.
Higher is better. The whisker spans each provider's best and worst category.
Linkup is absent: too little evidence to rank it in Search. It reappears — at the top — in the Extract lane below.
Quality is half the decision. An agent that waits nine seconds on a search feels broken even when the answer is good, and a provider that drops one call in twenty forces you to write retry logic. Both charts come from the same 216 searches the judges scored.
Solid bar is the median; the pale bar behind it is the slow 5% (p95). Lower is better.
Measured from a single European host. Absolute numbers shift with your region; the ordering travels.
Share of the 216 attempts that returned a usable response. Anything short of 100% is retry logic you have to write.
Up is better, left is faster. The corner you want is top-left.
Bubble size is cost per search; the free provider is drawn as a star.
No provider tops all nine search types, and in every single one of them the leader's confidence range still overlaps the runner-up's. That is the honest headline of this release.
Read these as leads, not wins: a leader here is the provider that happened to score highest, not one that has statistically separated from the pack.
Finding a page and reading it are different jobs. This lane hands each provider a URL and counts how many of the target facts survive the round trip.
Share of target facts recovered from the page. Higher is better.
Median time to fetch and return a page, 105 URLs. Pale bar is the slow 5%.
Dropped providers: loading…
The top of the table is not the cheapest, and the cheapest is not worthless. This is where you decide what the gap is worth to you.
Left is cheaper, up is better. The line traces the providers nothing else beats on both counts at once.
Scaled to spread cheap prices; ticks show actual USD.
Linkup isn't plotted: no rankable Search categories (insufficient evidence).
List prices (vendor pay-as-you-go); real bills vary with volume.
Provider identities were hidden from the judges, so the inputs carried anonymous codes instead of provider labels.
grok-4.5 and gpt-5.6-luna judged independently, so no single judge's quirks decide the outcome.
Every score comes with a confidence range; overlapping ranges are reported as statistical ties, not proven wins.