PA

LLM/agent evaluation stack: metered evaluator API

PATRONUS AIA3 · direct accessManage supplier listing

LLM/agent evaluation stack: metered evaluator API (Lynx hallucination detector, GLIDER reasoning judge, Percival agent-failure debugger) plus published eval datasets/benchmarks (FinanceBench 10,000 financial Q&A pairs, BLUR 573 tip-of-the-tongue pairs) and Digital World Models RL environments built on 1M+ world data artifacts from 5k+ expert contributors

A sample dataset retrieved from the supplier.

No sample attached yet

Request one and the supplier can attach it here.

Request sample

Universe, instruments and categories.

Categories
AI evaluation, guardrail and agent-training-environment platform (evaluator APIs + benchmark datasets)
Regions
US
Sample tickers
MSFTGOOGLAMZNNOWDDOGCRM

Patronus AI has not added their own details yet. Not yet on file:

  • Dataset
  • Data dictionary
  • Coverage
  • Provenance
  • Rights & data handling