That is a near-miss match: the entity was close enough to look right and wrong enough to hurt. Nothing crashed. The write, or the route, succeeded on the wrong row. This is a note about a failure mode teams are shipping into today — CRM, ERP, ticketing. It is not an Actuary case study. Keep your tracers and your eval gates. After ship you still need a quality floor on the match tasks you fear, not another green check beside “tool call completed”.
Signals
Why this hurts
Staff engineers and heads of AI feel this the moment an agent can resolve a person, an account, a ticket or a SKU and then act on what it found. Latency stays fine. Error rates stay fine. The span shows a lookup that returned and a write that went through. The damage is semantic: the row was real, and it was the wrong one.
Signal · usual read vs a near-miss match
| Signal | Usual read | Near-miss match |
|---|---|---|
| Search or retrieve | Returned a hit | The hit was the wrong twin |
| API status | Succeeded | Still green |
| Trace span | Tool ran | Looks healthy |
| CI evals | Candidate passed | Fixtures usually carry unique names |
| Human ticket | Someone noticed | How you find out |
Close cousins share surnames. Subsidiaries share domains. Tickets share subject lines. Fuzzy search exists to be helpful, and being helpful is exactly the condition a near-miss thrives in. Once the agent writes on the row it chose, you are in the territory of the silent wrong write.
Your stack
What teams do today (keep it)
Every part of this already has a good tool. None of them answers the question:
- Traces and gateways (Langfuse, Helicone) — a wrong but valid contact ID still reads as a success
- Offline evals and PR gates (Braintrust, LangSmith) — thin on twins, and blind to a provider change that never opens a pull request
- Human review on risky steps — real, and prone to rubber-stamping
- Confidence thresholds on the search — both candidates can clear the bar
- Spot audits of closed tickets — sparse, and late
Keep all of it. The pattern underneath is one sentence: a match is not the right match. The same argument about gates and what happens after them is CI evals green, production worse.
The floor
What to freeze
The question Actuary asks is narrow. On a frozen panel of match tasks, has served quality dropped below a floor you wrote down — with the false alarm budget fixed before you looked, and an alarm that still means something when you check it every day?
Three things to freeze:
- Twin panels: the pairs you already confuse, kept on purpose
- The task written as an outcome — pick THIS contact on THIS opportunity, not “a contact was found”
- A deterministic check wherever the right answer has an ID
Then the loop:
- Freeze the panel, the floor, and the false alarm budget
- Score the served route against that panel
- Watch it with anytime-valid evidence, so a daily glance invents nothing
- Send the alarm to Slack or email when the floor breaks
None of this sits inside the Salesforce transaction. It is a layer after ship: it does not block the write, it tells you when the match behaviour you were relying on has got worse. Watchtower is the public version of the instrument — live cells, the quiet ones published on purpose. The mathematics, and the regime where this detector is badly wrong, are on methodology.
Shapes
Shapes this takes
- Name twin — two contacts, one surname, one company
- Domain cousin — two subsidiaries in one holding group, one email domain
- Ticket subject collision — two open threads, the same subject line
- Stale ID in context — the agent reuses the ID it resolved three turns ago
- Top-1 overconfidence — fuzzy search always returns something
- A write on top of any of them — the near-miss becomes a silent wrong write
Boundary
Honest limits
Not claimed
Not a customer story. Not an argument that agents should stop resolving records. Not a replacement for your tracer, your eval suite, or the entity resolution and MDM work that keeps identities straight in the first place. Not SOC 2. Not a promise that every near-miss becomes a Watchtower incident — Study 001 is one trajectory. And a floor on match quality does not repair a knowledge base that is wrong underneath it.
Stack map
Where each instrument sits
Job · instrument · what to do about it
| Job | Instrument | What to do |
|---|---|---|
| Debug one bad run | Traces | Keep what you have |
| Block a bad candidate | CI evals | Keep |
| Cost and routing | Gateway | Keep |
| Risky mutations | Human review | Keep |
| Identity and dedupe | Entity resolution, MDM | Keep |
| A floor on served match quality, over time | Often missing | Where Actuary sits |
Next step
Next step
If the last CRM mess on your team was “almost the right contact”, and a human is the one who found it, that is the behaviour to put a floor under. Start two free Watch monitors and freeze a floor on that match behaviour.
Silent wrong write · CI evals green, production worse · how it works · the rest of the shelf.