It is easy to run the two together, because both draw graphs about models. They answer different questions. This note pulls them apart, so that you can keep Langfuse, Braintrust, Helicone or whatever you already run, and still be able to say whether quality has fallen below a line you wrote down in advance.
Observability
What observability covers
Tracing and gateway products record the operational path:
- Request path, prompt version, tool calls, and spans
- Latency, tokens, cost, and error rates
- Artifacts you open when a user reports a bad turn
That is the right instrument for request-level debug and for latency or cost SLOs. Langfuse, Helicone, and similar tools own that job. Actuary does not replace them. Product framing: actuary.sh; loop: how it works.
Observability alone does not provide:
- A quality floor frozen before the monitoring epoch started
- A false-alarm budget (α) that survives continuous peeking
- A design-bound claim that served quality stayed above a named floor
Those are monitoring and (when registered) certification concerns. Encoding them as ad-hoc dashboard thresholds usually recreates unbudgeted significance testing with better charts.
Offline evals
What offline evals cover
Eval studios and CI harnesses (Braintrust, LangSmith-style runners, custom suites) answer whether a candidate cleared a bar before ship. That is a gate. Gates matter. They operate on a different time slice than production:
Time slice · offline gate vs production monitor
| Concern | Offline CI eval | Production quality monitoring |
|---|---|---|
| When | Before or at deploy | After traffic / on a schedule |
| Unit | Candidate artifact | Served route over time |
| Peeking | Often one look per PR | Every day or every panel |
| Floor | Suite threshold | Floor + α frozen in the design |
| Failure mode | False green → ship a bad candidate | Silent drift → users notice first |
Green CI and worse production later is not a paradox. It is two instruments. See CI evals green, production worse.
The question
What quality monitoring answers
Actuary’s question is narrow:
Did production quality on a frozen panel drop below a pre-committed floor, with an anytime-valid alarm at a design-time false-alarm rate α?
Mechanics (how it works, methodology):
- Freeze panel composition, floor, and α before the first scored observation of an epoch.
- Score one complete panel — hosted probes across five task families with deterministic checkers, or scores you POST yourself.
- Update a betting e-process; alarm when E ≥ 1/α (default α = 0.05 → threshold 20).
- On registered designs, emit a signed offline-verifiable certificate. Descriptive streams stay labelled descriptive.
- False alarm budget
- α 0.05
- Alarm at
- E ≥ 20
The public surface is Watchtower: live model × family cells, including quiet spells. Incidents appear only after statistical and confound gates.
Stack map
Stack map (complement, not replacement)
Stack map · who owns which job
| Job | Typical owner | Actuary’s role |
|---|---|---|
| Traces / debug | Langfuse, OpenTelemetry | None — keep your tracer |
| Offline evals / PR gates | Braintrust, custom CI | Complement; keep the gate |
| Gateway / cost routing | Helicone, gateways | Complement |
| Frozen floor + anytime-valid alarm | Often missing | This layer |
| Offline-verifiable cert on registered design | Rare | Registered tier |
If the question is “Actuary vs Langfuse,” the accurate answer is with, not replace.
Peeking
Continuous peeking and naive quality alerts
Recomputing a fixed-horizon 0.05 test after every panel does not preserve a 5% false-alarm story. Anytime-valid e-processes keep a bound that holds at every look. Study 001 on Watchtower is one trajectory, not an operating-characteristic table. For FAR under dependence, use methodology.
Alert package
What belongs in the alert package
- Stream identity (suite / family / route)
- Floor, α, and current e-value (or log e)
- Onset diagnostic (validity attaches to the e-process, not the onset guess)
- Link to the incident trajectory and design hashes
When to add
When to add a quality monitor
Keep traces and CI. Add a quality monitor when silent degradation would hurt users, you can produce about one quality score per day (or use hosted probes), and you will freeze panel + floor + α and accept quiet periods.
Design-partner stage; do not claim SOC 2.
Next step
Next step
If you can see every span but still learn about quality from users, start a free Watch monitor.
Public record: Watchtower.