Actuary

Reference note

LLM quality monitoring is not tracing

Traces, latency charts and token dashboards tell you what a request did. That is observability, and you need it for debugging and for controlling spend. It is a different job from quality monitoring, which asks whether the answers are still good enough: a minimum score agreed before any data arrives, the same test cases every time, and an alarm that stays trustworthy even when you check it every day.

It is easy to run the two together, because both draw graphs about models. They answer different questions. This note pulls them apart, so that you can keep Langfuse, Braintrust, Helicone or whatever you already run, and still be able to say whether quality has fallen below a line you wrote down in advance.

Observability

What observability covers

Tracing and gateway products record the operational path:

  • Request path, prompt version, tool calls, and spans
  • Latency, tokens, cost, and error rates
  • Artifacts you open when a user reports a bad turn

That is the right instrument for request-level debug and for latency or cost SLOs. Langfuse, Helicone, and similar tools own that job. Actuary does not replace them. Product framing: actuary.sh; loop: how it works.

Observability alone does not provide:

  • A quality floor frozen before the monitoring epoch started
  • A false-alarm budget (α) that survives continuous peeking
  • A design-bound claim that served quality stayed above a named floor

Those are monitoring and (when registered) certification concerns. Encoding them as ad-hoc dashboard thresholds usually recreates unbudgeted significance testing with better charts.

Offline evals

What offline evals cover

Eval studios and CI harnesses (Braintrust, LangSmith-style runners, custom suites) answer whether a candidate cleared a bar before ship. That is a gate. Gates matter. They operate on a different time slice than production:

Time slice · offline gate vs production monitor

Concern Offline CI eval Production quality monitoring
When Before or at deploy After traffic / on a schedule
Unit Candidate artifact Served route over time
Peeking Often one look per PR Every day or every panel
Floor Suite threshold Floor + α frozen in the design
Failure mode False green → ship a bad candidate Silent drift → users notice first

Green CI and worse production later is not a paradox. It is two instruments. See CI evals green, production worse.

The question

What quality monitoring answers

Actuary’s question is narrow:

Did production quality on a frozen panel drop below a pre-committed floor, with an anytime-valid alarm at a design-time false-alarm rate α?

Mechanics (how it works, methodology):

  1. Freeze panel composition, floor, and α before the first scored observation of an epoch.
  2. Score one complete panel — hosted probes across five task families with deterministic checkers, or scores you POST yourself.
  3. Update a betting e-process; alarm when E ≥ 1/α (default α = 0.05 → threshold 20).
  4. On registered designs, emit a signed offline-verifiable certificate. Descriptive streams stay labelled descriptive.
False alarm budget
α 0.05
Alarm at
E ≥ 20

The public surface is Watchtower: live model × family cells, including quiet spells. Incidents appear only after statistical and confound gates.

Stack map

Stack map (complement, not replacement)

Stack map · who owns which job

Job Typical owner Actuary’s role
Traces / debug Langfuse, OpenTelemetry None — keep your tracer
Offline evals / PR gates Braintrust, custom CI Complement; keep the gate
Gateway / cost routing Helicone, gateways Complement
Frozen floor + anytime-valid alarm Often missing This layer
Offline-verifiable cert on registered design Rare Registered tier

If the question is “Actuary vs Langfuse,” the accurate answer is with, not replace.

Peeking

Continuous peeking and naive quality alerts

Recomputing a fixed-horizon 0.05 test after every panel does not preserve a 5% false-alarm story. Anytime-valid e-processes keep a bound that holds at every look. Study 001 on Watchtower is one trajectory, not an operating-characteristic table. For FAR under dependence, use methodology.

Alert package

What belongs in the alert package

  • Stream identity (suite / family / route)
  • Floor, α, and current e-value (or log e)
  • Onset diagnostic (validity attaches to the e-process, not the onset guess)
  • Link to the incident trajectory and design hashes

When to add

When to add a quality monitor

Keep traces and CI. Add a quality monitor when silent degradation would hurt users, you can produce about one quality score per day (or use hosted probes), and you will freeze panel + floor + α and accept quiet periods.

Design-partner stage; do not claim SOC 2.

Next step

Next step

If you can see every span but still learn about quality from users, start a free Watch monitor.

Start a free Watch monitor

Public record: Watchtower.