S

Scale

SCALEA3 · direct accessManage supplier listing

SWE-Bench Pro coding-agent benchmark: 1,865 tasks across 41 professional repos (public set = 731 GPL-repo instances, free via GitHub; private 276, held-out 858) + leaderboard; front for Scale's expert human-labeled training-data / RL-environment / eval business

A sample dataset retrieved from the supplier.

No sample attached yet

Request one and the supplier can attach it here.

Request sample

Universe, instruments and categories.

Categories
AI-agent evaluation benchmark dataset + leaderboard (from a training-data vendor)
Regions
US
Sample tickers
METANVDAMSFTGOOGLAMZN

Scale has not added their own details yet. Not yet on file:

  • Dataset
  • Data dictionary
  • Coverage
  • Provenance
  • Rights & data handling