Meet us atInvited Lecture @ National Artificial Intelligence Application Pilot Base·Sep 18, 2026·Hangzhou, ChinaAI Supply Chains Community Event·Sep 20, 2026·ShanghaiInvited Talk @ Compound Research Day: Data Factories·Sep 29, 2026·San Francisco, USANeurIPS 2026·Dec 6, 2026·Sydney, AustraliaInvited Talk @ NII Shonan·Mar 15, 2027·Hayama, Japan
Meet us atInvited Lecture @ National Artificial Intelligence Application Pilot Base·Sep 18, 2026·Hangzhou, ChinaAI Supply Chains Community Event·Sep 20, 2026·ShanghaiInvited Talk @ Compound Research Day: Data Factories·Sep 29, 2026·San Francisco, USANeurIPS 2026·Dec 6, 2026·Sydney, AustraliaInvited Talk @ NII Shonan·Mar 15, 2027·Hayama, Japan

The numbers on our landing page

Our landing page leads with three numbers: 45.6 minutes, $3.58, and 88.7%. This post states exactly what each one measures, how we measured it, what we have not measured yet, and one framing we tested and dropped after our own blind-adjudicated pilot went against us.

Luis Oala
The numbers on our landing page

Our landing page now leads with three numbers: 45.6 minutes, $3.58, and 88.7%. This post states exactly what each one measures, how we measured it, what we have not measured yet, and one framing we tested and dropped after our own pilot went against us.

Why we benchmark ourselves

Our thesis is that the data most likely to move a model or a strategy is not listed in any catalog, and that the binding constraint is access, not model capability. A company built on that claim owes you measurements of its own system that it would demand from anyone else's.

What we measured, and how

Three measurements, all of our own production system. None of them is a head-to-head comparison, and we say so plainly: there is no incumbent on the other side of any of these numbers. Every precise published figure we could find for the cost and duration of a traditional data deal traced back to a single citation we could not locate, so we do not use them. The defensible statement is narrative: the standard process runs weeks to months. Until we measure an incumbent baseline ourselves, no multiplier appears on our page.

Latency. We queried our production request log for every completed request: n = 207. For each, we took the wall-clock time from submission to completion of the full pipeline, including source discovery and enrichment.

Cost. We ran cost accounting over 110 delivered production tables. The figure is model spend only, estimated from CLI logs. It is a floor: it excludes infrastructure and third-party enrichment credits.

Catalog absence. We took 565 delivered sources that had passed deep verification and looked each one up in 14 indexed data catalogs, including Datarade, AWS Data Exchange, Snowflake, and Neudata.

The three numbers

45.6 minutes is the median time to complete a request, over the 207 completed production requests. The P95 is 3.65 hours: slow requests take hours, not days, and the median covers only requests that completed. We also keep a re-runnable harness on 50 machine-fetchable endpoints from our registry — the easiest access tier — where the request-to-verified-sample roundtrip has a median of about 11 minutes. The production number is four times slower, and it is the one on the page.

$3.58 is the median LLM cost to produce one delivered table (n = 110). Per delivered contact row, it is $0.097. Again: a floor, model spend only.

88.7% is the share of the 565 deep-verified delivered sources for which we could locate no ready-to-buy product in any of the 14 catalogs. The denominator is self-selected: it is our own delivered slate, and that slate skews to the frontier by design. Read it as a description of what we deliver, not a measurement of the whole market.

One more measurement did not fit on the page. Of the 27,899 claims we have delivered, 93.3% carry a per-claim evidence record, and every delivered row traces to pages fetched inside its request window. On the newest delivered table, all 51 page fetches and all 857 claims are timestamped inside the 54-minute request window — a single table, so an existence proof, not a distribution.

A framing we tested and dropped

Earlier internal drafts framed us as better at contact-name discovery. We tested it: a blind-adjudicated pilot of 30 supplier cases, our pipeline against a web-only frontier LLM. On the outcome a buyer pays for — the correct current decision-maker plus a usable address — we led 14/30 to 6/30 (p ≈ 0.039). On bare names, the frontier model beat us 83% to 53%, and 13.2% of our name claims failed adjudication, against 0% of theirs. A 30-case pilot is not a verdict either way. We dropped the framing, and a 150-case rerun is scheduled. We publish the facets that went against us for the same reason we publish every n: the numbers we keep mean little if you cannot see the ones we discard.

What we have not measured yet

  • An incumbent baseline: the cost and duration of traditional procurement, measured rather than cited.
  • Catalog absence on a slate we did not pick: the same 14-catalog audit re-run on an externally chosen set of sources.
  • The 150-case contact rerun.
  • Fully loaded cost per table, beyond model spend.

The public benchmark

Our public benchmark work asks the same question from the other side. The Information Frontier Benchmark (May 2026) ran 1,123 provisioning attempts across access tiers. Frontier models provisioned 80% of open-web sources and 0% of sales-gated ones — an 80-point spread across access tiers, against a 7-point spread across models (29, 23, and 22 on the same tasks). Pooled across the gated tiers, 10.8% of attempts provisioned, and payment walls were uncrossable by construction: the benchmark spends $0. The barrier is access, not capability. That finding is why this product exists, and why we measure it the way we do.

For the worldview behind the benchmark, see The Information Frontier.

Expand your information frontier.

Get Started