S
Scale
SCALEA3 · direct accessManage supplier listing
SWE-Bench Pro coding-agent benchmark: 1,865 tasks across 41 professional repos (public set = 731 GPL-repo instances, free via GitHub; private 276, held-out 858) + leaderboard; front for Scale's expert human-labeled training-data / RL-environment / eval business
Sample
A sample dataset retrieved from the supplier.
No sample attached yet
Request one and the supplier can attach it here.
Request sampleCoverage
Universe, instruments and categories.
- Categories
- AI-agent evaluation benchmark dataset + leaderboard (from a training-data vendor)
- Regions
- US
- Sample tickers
- METANVDAMSFTGOOGLAMZN
Scale has not added their own details yet. Not yet on file:
- Dataset
- Data dictionary
- Coverage
- Provenance
- Rights & data handling