Encyclopedia Evalica / Deployment / Caching

Caching illustration

Caching

/'ka.shihng/The process of storing and reusing LLM responses for identical inputs, reducing cost and latency. Caching is especially impactful for repeated prompts and common workflows. (noun)

Caching cut our average latency in half for repeat questions.

Customer example

Graphite used Braintrust prompt caching to manage costs while benchmarking models for its AI code reviewer. Read more

Related Deployment terms

From the docs

Get started with Evals

Braintrust is the observability platform for agents in production. By actively applying intelligence to agent traces and automatically surfacing critical patterns, Braintrust helps teams at Notion, Stripe, Box, OpenAI, and Cloudflare ship quality agents at scale.

Start building