LIVE

Rheostat Summit 2026 — twelve talks on shipping agents safely. Oct 14, San Francisco.

Save your seat

Rheostat

0 Book a demo
Experiments

Ship the variant
that won.

Every flag is already an experiment. Pick a metric, split the traffic and let the stats engine tell you when it is safe to ship.

checkout-cta14 days · 182,400 usersReady to ship
A · control"Continue"3.12%
B · variant"Pay £42 now"3.25%
Uplift
+4.2%
Chance to beat A
97.8%
Revenue / user
+£0.31
Ship variant B to 100%

From hypothesis to shipped
in five steps.

  1. 01HypothesisWrite what you expect and which metric proves it.
  2. 02SplitPick a flag and a percentage. No new deploy.
  3. 03GuardrailsAuto-stop if latency or errors move the wrong way.
  4. 04ReadSequential stats: peek any time without p-hacking.
  5. 05ShipOne click moves the winner to 100% and archives the loser.

Last week’s results.

pricing-page-annualWinner

+11.3%annual plan take-rate

onboarding-checklistNo effect

+0.4%activation in 7 days

agent-summary-v2Loser

−6.1%tickets resolved by agent

Metrics you
already trust.

Pull metrics from your warehouse or define them in SQL. Every experiment, flag and agent run can use the same definitions.

  • conversion_rateunit · ratio
  • revenue_per_userunit · £
  • p95_latencyunit · ms
  • churn_30dunit · %
  • activation_7dunit · %
  • tokens_per_sessionunit · count
  • support_ticketsunit · count
  • refund_rateunit · %
  • agent_resolutionunit · %

“We stopped arguing about which idea was better. Now we test it on Tuesday and ship the winner on Friday.”

Jonah OseiVP Product, Brightline
62
experiments shipped by Brightline last quarter, up from 9

Experiment questions.

Which statistics do you use?
Sequential testing with always-valid confidence intervals, so you can look at results every day without inflating errors.
How long should a test run?
Until the result is conclusive or the minimum effect is ruled out. Most tests finish in one to two weeks.
Can I test prompts and models?
Yes. Any flag can be an experiment: a prompt, a model, a ranking function or a button colour.
What are guardrail metrics?
Metrics that must not get worse, like latency or refunds. A breach stops the test automatically.

Stop guessing,
start measuring.

Turn any flag into an experiment and get a plain-English verdict when it is done.