Encyclopedia Evalica / Datasets / Expected output (ground truth)

Expected output (ground truth)
/ih'kspeh.ktuhd 'ow.tpuut grownd trooth/The reference answer/behavior a dataset record defines as correct (when doing reference-based evals). Ground truth can be human-written, programmatically derived, or imported from production. (noun)
“The expected output includes the exact policy language and a citation.”
Related Datasets terms
From the docs
Get started with Evals
Braintrust is the observability platform for agents in production. By actively applying intelligence to agent traces and automatically surfacing critical patterns, Braintrust helps teams at Notion, Stripe, Box, OpenAI, and Cloudflare ship quality agents at scale.
Start building