That is the silent wrong write: an agent changes a system of record, the tooling reports success, and the quality problem only surfaces when a person notices. This is a note about a failure mode teams are shipping into today — CRM, ERP, ticketing. It is not an Actuary case study. Keep your traces and your eval gates. After ship you still need a floor on the behaviour you fear, not another tick beside “tool call completed”.
Signals
Why this hurts
Latency alerts fire when the stack is slow. Error rates fire when it throws. Traces show you the tool call. A silent wrong write clears all three bars.
Signal · usual read vs a silent wrong write
| Signal | Usual read | Silent wrong write |
|---|---|---|
| API status | Succeeded | Still says succeeded |
| Trace span | Tool ran | Looks healthy |
| CI evals | Candidate passed | Often still passing |
| Human ticket | Someone noticed | How you find out |
The damage is delayed. Pipeline numbers lie for a quarter. A customer gets the wrong email. An account executive spends a morning untangling merged records. Staff engineers and heads of AI feel this the moment the agent can write, not only draft.
Your stack
What people already do (keep it)
Teams already buy good tools for parts of this:
- Tracing (Langfuse, LangSmith) — see the tool call and its arguments
- Eval studios and gates (Braintrust) — offline suites before deploy
- Gateways (Helicone) — cost, routing, request logs
Keep them. Actuary does not replace that stack; monitoring sits beside tracing, and that difference is the subject of monitoring is not tracing.
The gap is narrow and specific: “the tool call succeeded” is not “the write was correct for the business”. Online judges and dashboards help. But if you check them every day without a line written down in advance and a budget for false alarms, the thresholds drift until someone mutes the channel.
Two instruments
Two instruments, two questions
- Ship gate: is this candidate good enough to deploy?
- Served-route floor: is production still above the line we wrote down?
Providers change. Prompts drift. Aggregator paths shift. None of that opens a pull request, which is why a green gate and a worse production system are not a contradiction. See CI evals green, production worse.
The floor
What a quality floor looks like here
Before you look at a single new score, write three things down:
- The test cases that encode the writes you fear — wrong owner, wrong amount, wrong entity match, whatever you can score
- The lowest score you will accept before you call it too wrong
- How often you are willing to cry wolf — say, at most one false alarm in twenty
Then watch with an alarm that stays honest at every look, so checking daily does not invent a result. The mathematics, and the regime where this detector is badly wrong, are on methodology.
A public example of the instrument — not a CRM case study — is Watchtower: live cells including the quiet ones.
Boundary
What this note is not
Not claimed
Not a claim that Actuary watched your Salesforce. Not a customer story. Not “replace your tracer”. Not SOC 2. Not a promise that every silent wrong write can be scored.
Next step
Next step
If an agent can change a system of record, pick the one failure you fear most, write the floor down, and watch it after ship.