Encyclopedia Evalica / Datasets / Expected output (ground truth)

Expected output (ground truth) illustration

Expected output (ground truth)

/ih'kspeh.ktuhd 'ow.tpuut grownd trooth/The reference answer/behavior a dataset record defines as correct (when doing reference-based evals). Ground truth can be human-written, programmatically derived, or imported from production. (noun)

The expected output includes the exact policy language and a citation.

Related Datasets terms

From the docs

Get started with Evals

Braintrust is the observability platform for agents in production. By actively applying intelligence to agent traces and automatically surfacing critical patterns, Braintrust helps teams at Notion, Stripe, Box, OpenAI, and Cloudflare ship quality agents at scale.

Start building