Rheostat

Book a demo
Guides

Progressive delivery for LLM prompts

A prompt change is a production change. It alters what your product says to every customer, and it can fail in ways no unit test will catch.

Why prompts are deploys

A new prompt ships to every conversation at once unless something stops it. That is the same risk as a code deploy, with less tooling around it.

Version every prompt

Store prompts next to the code that calls them, give each a semantic version, and reference the version from a flag rather than hard-coding the text. The flag becomes the switch between v14 and v15 — and the audit log records who flipped it.

const prompt = await rh.flag("support-prompt", user)
// → "v15" for 5% of users, "v14" for everyone else
const reply = await model.run(prompts[prompt], ticket)

Roll out by percentage

Start at 5%, hold for a day, and let the dial move only while the numbers stay flat. Every step is a chance to stop.

Roll back on a judge

Percentages only help if something is watching. Attach a judge — PII, tone, accuracy — and a rule that dials the prompt to 0% when the score drops. Ours fired twice in the first week. Nobody noticed except the audit log.

The first rollback happened at 2 a.m. and took 900 milliseconds. We read about it over breakfast.

What we learned

Note

Judges add latency to the eval path only when sampled. Start at 10% of traffic and raise it once you trust the scores.