Nothing here says you should throw away your eval harness. It explains why alert fatigue is usually a statistics problem rather than a discipline problem, how we spend a fixed false-alarm budget instead of gambling it away on every look, and where our own instrument stops working. The limits are set out on the methodology page.
False alarm
What “false alarm” means here
A false alarm is an alert that fires while quality is still at or above the floor you committed to. Three mistakes:
- Unbudgeted looks. Recompute a p < 0.05 test after every panel.
- Post-hoc floors. Choose the threshold after seeing the dip.
- Dependence denial. Shared latent shocks inflate empirical FAR above nominal α.
Actuary sets α at design time and alarms when e-value crosses 1/α. Default α = 0.05 → threshold 20. That is a budget, not a promise of zero false alarms.
- False alarm budget
- α 0.05
- Alarm at
- E ≥ 20
Peeking
Continuous peeking is the production default
Offline CI often gets one look per PR. Production monitors get a look every day. That breaks fixed-horizon error control if each look is treated as the only look. Methodology contrasts fixed-horizon-style p-values vs e-process on the 2026-07-24 degraded lane.
Independence
Independence is a premise
Anytime-valid guarantees rely on a conditional-mean / independence-style premise. If wrong, a certificate still attests computation under the design — it does not prove the premise from the data.
56.38%
Measured FAR under block-5 dependence
Published stress: under temporally dependent scores — one latent component shared across 5 consecutive observations — measured FAR is 56.38% against nominal α = 0.05, over 10,000 seeded runs. Methodology — where it fails.
Cross-stream multiplicity (daily e-BH) is a separate gate.
Design-time α
What design-time α buys
- Quiet periods expected
- Alarms latch against the day-zero threshold
- Changing panel, floor or α opens a new epoch
Simulation, cited in context and not as live model measurement: the null arm returned a 0.252% false-alarm rate over a 20-panel horizon — methodology.
Study 001
Study 001 is one trajectory
openai/gpt-4.1-nano, frozen 5-item math panel, floor 0.70, α = 0.05.
Study 001 · two lanes · one trajectory each
| Lane | Panels | Outcome | Evidence | Bound |
|---|---|---|---|---|
| Stable | 20 | No alarm | 0.79 | — |
| Degraded | 12 | Alarm at block 12 | 29.96 | ≤ 19 |
Not a frequency claim or customer case study.
Tightening
Why tightening the threshold often makes fatigue worse
Raising the bar until Slack goes quiet silently spends undeclared α or couples the floor to last week’s noise. Declare α, freeze the panel, treat muting as a design change.
Alongside
How this sits next to existing tools
Langfuse/Braintrust/Helicone do not provide a frozen quality floor with anytime-valid peeking control. Put Actuary beside them. Do not rip them out.
Checklist
Checklist before trusting an LLM quality alert
- Floor and α before epoch
- Panel hash-committed
- Alarm valid under continuous peeking
- Independence/dependence stated
- Multiplicity handled
- Quiet publishable
If fewer than four hold, the alert is a heuristic.
Next step
Next step
Watchtower: the public record. Start a monitor: signup.