A prompt change is a production change. It alters what your product says to every customer, and it can fail in ways no unit test will catch.
Why prompts are deploys
A new prompt ships to every conversation at once unless something stops it. That is the same risk as a code deploy, with less tooling around it.
Version every prompt
Store prompts next to the code that calls them, give each a semantic version, and reference the version from a flag rather than hard-coding the text. The flag becomes the switch between v14 and v15 — and the audit log records who flipped it.
const prompt = await rh.flag("support-prompt", user)
// → "v15" for 5% of users, "v14" for everyone else
const reply = await model.run(prompts[prompt], ticket)
Roll out by percentage
Start at 5%, hold for a day, and let the dial move only while the numbers stay flat. Every step is a chance to stop.
Roll back on a judge
Percentages only help if something is watching. Attach a judge — PII, tone, accuracy — and a rule that dials the prompt to 0% when the score drops. Ours fired twice in the first week. Nobody noticed except the audit log.
The first rollback happened at 2 a.m. and took 900 milliseconds. We read about it over breakfast.
What we learned
Note
Judges add latency to the eval path only when sampled. Start at 10% of traffic and raise it once you trust the scores.