PA
LLM/agent evaluation stack: metered evaluator API
PATRONUS AIA3 · direct accessManage supplier listing
LLM/agent evaluation stack: metered evaluator API (Lynx hallucination detector, GLIDER reasoning judge, Percival agent-failure debugger) plus published eval datasets/benchmarks (FinanceBench 10,000 financial Q&A pairs, BLUR 573 tip-of-the-tongue pairs) and Digital World Models RL environments built on 1M+ world data artifacts from 5k+ expert contributors
Sample
A sample dataset retrieved from the supplier.
No sample attached yet
Request one and the supplier can attach it here.
Request sampleCoverage
Universe, instruments and categories.
- Categories
- AI evaluation, guardrail and agent-training-environment platform (evaluator APIs + benchmark datasets)
- Regions
- US
- Sample tickers
- MSFTGOOGLAMZNNOWDDOGCRM
Patronus AI has not added their own details yet. Not yet on file:
- Dataset
- Data dictionary
- Coverage
- Provenance
- Rights & data handling