Encyclopedia Evalica / Datasets / Dataset

Dataset
/'day.tuh.seht/A versioned collection of test cases used to run evals and track improvements over time. Versioning matters because it keeps comparisons reproducible. (noun)
“We added 50 new examples to the dataset and re-ran the eval.”
Related Datasets terms
From the docs
Get started with Evals
Braintrust is the observability platform for agents in production. By actively applying intelligence to agent traces and automatically surfacing critical patterns, Braintrust helps teams at Notion, Stripe, Box, OpenAI, and Cloudflare ship quality agents at scale.
Start building