Actuary

Reference note

LLM false alarm rate under continuous peeking

Most LLM quality alerts are a threshold on a dashboard with a webhook behind it. They feel rigorous right up to the day somebody mutes the channel. Two things are usually missing: a stated false alarm rate that survives you checking the dashboard whenever you like, and an honest note about what that rate assumes.

Nothing here says you should throw away your eval harness. It explains why alert fatigue is usually a statistics problem rather than a discipline problem, how we spend a fixed false-alarm budget instead of gambling it away on every look, and where our own instrument stops working. The limits are set out on the methodology page.

False alarm

What “false alarm” means here

A false alarm is an alert that fires while quality is still at or above the floor you committed to. Three mistakes:

  1. Unbudgeted looks. Recompute a p < 0.05 test after every panel.
  2. Post-hoc floors. Choose the threshold after seeing the dip.
  3. Dependence denial. Shared latent shocks inflate empirical FAR above nominal α.

Actuary sets α at design time and alarms when e-value crosses 1/α. Default α = 0.05 → threshold 20. That is a budget, not a promise of zero false alarms.

False alarm budget
α 0.05
Alarm at
E ≥ 20

Peeking

Continuous peeking is the production default

Offline CI often gets one look per PR. Production monitors get a look every day. That breaks fixed-horizon error control if each look is treated as the only look. Methodology contrasts fixed-horizon-style p-values vs e-process on the 2026-07-24 degraded lane.

Independence

Independence is a premise

Anytime-valid guarantees rely on a conditional-mean / independence-style premise. If wrong, a certificate still attests computation under the design — it does not prove the premise from the data.

56.38%

Measured FAR under block-5 dependence

Published stress: under temporally dependent scores — one latent component shared across 5 consecutive observations — measured FAR is 56.38% against nominal α = 0.05, over 10,000 seeded runs. Methodology — where it fails.

Cross-stream multiplicity (daily e-BH) is a separate gate.

Design-time α

What design-time α buys

  • Quiet periods expected
  • Alarms latch against the day-zero threshold
  • Changing panel, floor or α opens a new epoch

Simulation, cited in context and not as live model measurement: the null arm returned a 0.252% false-alarm rate over a 20-panel horizon — methodology.

Study 001

Study 001 is one trajectory

openai/gpt-4.1-nano, frozen 5-item math panel, floor 0.70, α = 0.05.

Study 001 · two lanes · one trajectory each

Lane Panels Outcome Evidence Bound
Stable 20 No alarm 0.79 —
Degraded 12 Alarm at block 12 29.96 ≤ 19

Not a frequency claim or customer case study.

Tightening

Why tightening the threshold often makes fatigue worse

Raising the bar until Slack goes quiet silently spends undeclared α or couples the floor to last week’s noise. Declare α, freeze the panel, treat muting as a design change.

Alongside

How this sits next to existing tools

Langfuse/Braintrust/Helicone do not provide a frozen quality floor with anytime-valid peeking control. Put Actuary beside them. Do not rip them out.

Checklist

Checklist before trusting an LLM quality alert

  • Floor and α before epoch
  • Panel hash-committed
  • Alarm valid under continuous peeking
  • Independence/dependence stated
  • Multiplicity handled
  • Quiet publishable

If fewer than four hold, the alert is a heuristic.

Next step

Next step

Read the methodology

Watchtower: the public record. Start a monitor: signup.