Meet us atInvited Lecture @ National Artificial Intelligence Application Pilot Base·Sep 18, 2026·Hangzhou, ChinaAI Supply Chains Community Event·Sep 20, 2026·ShanghaiInvited Talk @ Compound Research Day: Data Factories·Sep 29, 2026·San Francisco, USANeurIPS 2026·Dec 6, 2026·Sydney, AustraliaInvited Talk @ NII Shonan·Mar 15, 2027·Hayama, Japan
Meet us atInvited Lecture @ National Artificial Intelligence Application Pilot Base·Sep 18, 2026·Hangzhou, ChinaAI Supply Chains Community Event·Sep 20, 2026·ShanghaiInvited Talk @ Compound Research Day: Data Factories·Sep 29, 2026·San Francisco, USANeurIPS 2026·Dec 6, 2026·Sydney, AustraliaInvited Talk @ NII Shonan·Mar 15, 2027·Hayama, Japan

Blog

Insights from the data frontier

On data infrastructure, procurement, and the evolving data economy.

Latest Posts

Changelog #5: Public Index, Signals, and Sequences — Completing the Roundtrip on Automated Data Procurement

Alt-data's listed universe is now public and available on Brickroad. The Public Index maps the majority of suppliers listing on the most popular data directories, like Neudata, Eagle Alpha, Datarade, AWS Data Exchange, and Snowflake Marketplace. Signals ships with a daily feed of frontier sources tied to trading and trending news. Our Access Agents now pull samples from open paths. Outbox Sequences and Templates allow you to run prospecting and outreach automatically. Ground your data licenses against real usage. And first-time suppliers can now get license-ready with a checklist we built to spec with real buyers.

Freeman Lewin

Popular Data Vendors August 2026

List of popular neo-vendor endpoints in August 2026 across the Brickroad data frontier registry

Luis Oala

Cracking the information frontier

The information frontier grows faster than today's data access protocols, yet the most valuable computations rely on exactly that frontier data. We are building the data multiplexer, a technology that anyone can point at the information frontier to discover, access, route and manage data flow for their most valuable computations. The reference user study below spanning usage from May to August 2026 illustrates the gains and scale of our technology over existing data procurement protocols.

Luis Oala

The numbers on our landing page

Our landing page leads with three numbers: 45.6 minutes, $3.58, and 88.7%. This post states exactly what each one measures, how we measured it, what we have not measured yet, and one framing we tested and dropped after our own blind-adjudicated pilot went against us.

Luis Oala

Changelog #4: Sources, Agent Outboxes, and API Access — From Discovery to Acquisition

The other half of the sourcing marathon is live on Brickroad. Sources is your organization's system of record for every supplier your agents surface. Agent Outboxes handle approval-gated, pseudonymous outreach and sample collection. And our first public API puts discovery in your scripts, agents, and internal tooling — not just the dashboard.

Freeman Lewin

Changelog #3: Source Streams — Data Supplier Discovery on Autopilot

Source Streaming is now live on Brickroad. Set your thesis once, and your agent runs continuously, notifying you the moment a new data supplier comes online. Plus new APAC supplier discovery across China, Japan, Korea, and India; established-supplier views alongside frontier sources; and workspaces with category views for organizations running concurrent queries.

Freeman Lewin

Neither Alone, Both in Sequence: Human-Agent Collaboration, Intellect, and the Information Frontier

Part 1 of The Data Multiplexer Series. The information frontier is structural and ever-widening; reaching it requires not just speed but intellect. Defining intellect, and why the frontier is reached by humans and agents only when paired in deliberate sequence — neither alone.

Freeman Lewin

Croissant Tasks: Machine-Actionable Metadata for Reproducible ML Evaluations

Croissant Tasks is a declarative metadata format that turns benchmarks and competitions into machine-actionable specifications. It enables conceptual reproducibility: verifying a scientific claim through an independently generated implementation rather than brittle source-code replication.

Luis Oala

Making the Discrete Continuous: Synthetic RAW Augmentations for Low-Light Person Detection

Real datasets are sparse and uneven, which makes it hard to evaluate vision models where it matters most. By synthesizing physically faithful low-light RAW samples, we can turn a discrete, long-tailed variable into a continuous, controllable one and fairly characterize pedestrian detection in the dark.

Luis Oala

Croissant Baker: Local-First Metadata Generation for Governed ML Datasets

Croissant has become the metadata standard for ML datasets, but generating it usually means uploading data to a public platform — impossible for clinical, government, and enterprise data. Croissant Baker generates validated Croissant metadata locally, directly from a dataset directory, reaching 97-100% agreement with ground truth across domains and scaling to MIMIC-IV's 886 million rows.

Luis Oala

The Information Frontier

A reductionist view of machine learning as a perpetual data refinery, and a re-calibration of its primitives. Why the information frontier is perpetually expanding, what physics says about ever collapsing it, and what it implies for the learning systems we build and study.

Luis Oala

The Geometry of Data Markets

Why data marketplaces won't lead to data liquidity: lessons from history. Data liquidity is real, but it requires the right shape. Not a catalog. Not a directory. Not a platform. A multiplexer, with agents underneath, routing the right data to the right endpoint at the right time.

Freeman Lewin

OpenML: Insights from 10 Years and More Than a Thousand Papers

A decade of OpenML, the open-source platform that turns machine-learning experiments into open, linked, and reusable knowledge. We look at the state of the ecosystem, how community-curated datasets, tasks, and benchmark suites have powered 1,500+ studies, and the lessons learned from building open-science infrastructure for ML.

Luis Oala

Croissant: A Metadata Format for ML-Ready Datasets

Working with data is still a key friction point in machine learning. Croissant is a metadata format that creates a shared representation across ML tools, frameworks, and platforms — making datasets discoverable, portable, and interoperable. It is already supported across repositories spanning hundreds of thousands of datasets.

Luis Oala

DMLR: Data-Centric Machine Learning Research — Past, Present and Future

Drawing on discussions at the inaugural DMLR workshop at ICML 2023, this editorial outlines why community engagement and infrastructure are essential to creating the next generation of public datasets — and charts a collective path to sustain them for scientific, societal, and business impact.

Luis Oala

Expand your information frontier.

Get Started