Actuary

Model behaviour, on the record

Know the day your AI gets worse.

Evals tell you it passed yesterday. Actuary watches production — when quality drops, you get an alarm you can defend and a signed certificate anyone can recheck.

Start free

Two monitors free, forever · we set the first one up with you

One monitor, one task. This demo watches an invoice-extraction agent against a floor frozen in advance.

MONITOR · INVOICE-EXTRACTION AGENT 5-ITEM PANEL · FLOOR 0.70 · α 0.05 STABLE
QUALITY · PANEL MEANFLOOR1.000.900.800.600.70EVIDENCE AGAINST “THE FLOOR HOLDS”ALARM · 1/α205210.78CHECKS →
How long has it watched? 20 Complete panels · one per check
Is quality holding? 0.902 Panel quality above the 0.70 floor
Should anyone be paged? 0.86 of 20 Evidence against “the floor holds” the bet has not paid off

The estimator runs in your browser. Inject a fault and watch the evidence answer.

When this fires

One alert in Slack and email · an incident on the record · a signed certificate on registered designs

How it breaks in production — try one

Patient on purpose — evidence compounds only while quality sits below the floor

SELF-CHECK AT LOAD · ENABLE JAVASCRIPT TO RECOMPUTE CERTIFICATE 001 HERE

The estimator that reproduced the 2026-07-24 certificate, bit for bit. Faults are simulated; the mathematics is not. After a fault, time compresses.

#eng-alerts 3 members
Actuary APP 09:14

The Actuary detected a threshold crossing for tenant northwind — extraction quality fell below the frozen floor.

Model gpt-5-mini
Task family extraction
log10 e-value 1.476582
Onset estimate observation 8 at 2026-07-31 09:06:12 UTC (estimate)
Alert ID 0b41e2c7-9f3a-4d21-b6ce-8a5f21c704d9
Stream 4c7d1f60-2b58-4a0e-9c3d-1d6f0b7a5e42
Incident https://actuary.sh/v1/incidents/9d2b8f14-6c07-4e3a-8b51-0f7c2a94d6b3
actuary.sh

Start free this alarm, on your task, from one score a day

Sept 17, 2025 · Anthropic engineering postmortem

“The evaluations we ran simply didn’t capture the degradation users were reporting… we lacked a clear way to connect these to each of our recent changes.”

Weights get pinned. Served behaviour does not.

The provider needed six weeks to see it inside their own system — you are one call downstream, with no instrument at all.

What you get

A record you can hand to someone else.

Every monitor keeps one: frozen panels, a running interval on quality, the whole epoch lineage — and no way to quietly edit any of it.

SUPPORT-AGENT · gpt-5-mini · EPOCH 3 · PANEL FROZEN BEFORE DATA · LINEAGE v1 → v2 → v3 SYNTHETIC
0.600.700.800.901.00PANEL QUALITY · ONE FROZEN PANEL PER DAYFLOOR 0.70DAYS → SEGMENT HEIGHT = CONFIDENCE INTERVALCONFIDENCE SEQUENCE · VALID AT EVERY LOOK
ALERT ROUTING · SLACK + EMAIL — ARMED · NO-DATA DAYS · RESERVED SWATCH, NEVER GREEN · ONE SEGMENT PER DAY — POSITION IS THE PANEL MEAN, HEIGHT THE CONFIDENCE INTERVAL · THE WATCHTOWER SERVES THE SAME PAGE FOR THE MODELS WE WATCH OURSELVES →

If you ship an LLM feature to real users and can send one quality score a day, you’re a fit. That is the whole bar.

POST /v1/scores
{"suite":"support-agent","suite_version":3,
 "model":"gpt-5-mini","score":0.86,
 "ts":"2026-07-29T09:14:07Z"}

One JSON call for the whole evidence machinery — alarms, history, the API. A certificate needs a design frozen with you first.

The guarantee

An alarm that says what it is worth.

If quality never fell below your floor, the chance of the evidence ever reaching the alarm level is at most 1 in 20 — no matter how often you look. Ordinary monitoring taxes every look with false alarms until you ignore it. This one already paid. When it fires, you act.

An eval suite

tells you a score, today.

A drift score

tells you something moved.

An Actuary alarm

tells you what the evidence is worth, at every look.

1 in 20

The false-alarm bound · at every look · for the life of one frozen design

Under the independence premise — see where it fails ↓

Proof

We made the instrument break on purpose first.

We registered an alarm before asking anyone to trust one: one lane served normally, one degraded on purpose, the bound written down first.

FIG. 002 — THE BREAK REGISTERED STUDY 001 · 2026-07-24
EVIDENCE · e-VALUE, LOG SCALE0.51520ALARM THRESHOLD · 1/α = 20PARITY · THE BET HAS PAID NOTHINGBOUND ≤ 19ALARM · PANEL 12 · e 29.96DEGRADED LANESTABLE LANE · CLOSES 0.79PANEL 15101520EVIDENCE · e-VALUE, LOG SCALE120ALARM · PANEL 12e 29.96DEGRADED LANESTABLE LANE · CLOSES 0.79PANEL 114
Stable lane0.79
Alarm panel12
Bound · blocks from onset≤ 19
Degraded close29.96
Prospective validation, 2026-07-24 · registered study 001 · openai/gpt-4.1-nano via OpenRouter · one frozen 5-item maths panel · floor 0.70 · α 0.05 · one real trajectory, not a frequency claim · fixed-horizon comparison on the methodology page.

Across all twenty stable panels the evidence never rose above its starting point — max log e is 0 on the signed certificate.

ACTUARY · REGISTERED STUDY 001 · PROSPECTIVE

Certificate of monitoring

actuary-certificate-v4 · the actual artifact,
rendered — recompute it
Run
real-validation-prospective-fixed-panel-2026-07-24
Subject
openai/gpt-4.1-nano · via OpenRouter
Design frozen
2026-07-24 · 5-item maths panel · floor 0.70 · α 0.05
Registered bound
alarm within 19 blocks of onset · pre-registered separately
Result
alarm at panel 12 · log e 3.3999562471954308
Period
2026-07-24 17:19:34 → 17:24:10 UTC
Payload hash
35aefb0e387dd7f1784c8499d4dff5f99d45d562b926a539a3b4ee7fc08a58c8
Signed
Ed25519 · key 079ff1a1fc426e9570f896537edf2fe5612768f5e5ff8e05b1cbd2b299ca1f25
Chain
previous payload hash null — genesis
Fields read from certificate.json · the bound was registered separately, before the run Signing-key fingerprint · one cell per hex digit

DESIGN PARTNERS · REGISTERED DESIGNS

Hosted registered checks

We freeze one task into a registered suite — panel, floor, α — before the data arrives. Every alarm ships a certificate.

$ actuary verify certificate.json
PASS canonical form
PASS signature · Ed25519, payload hash valid
PASS hash chain · certificate is genesis

No database and no private key — the verifier replays every block

The public record

We run it on public models, in public.

Our own monitors, pointed at hosted models, published unedited — alarms, quiet spells, mistakes. Nothing can be backfilled, not even by us.

Open the Watchtower Public · append-only · citable — the record opens the day the clock starts

Where it fails

We publish the conditions under which our own detector misfires.

Read this as an assumption-sensitivity result for the current estimator — not the product's rate in service.

56.38% False-alarm rate under block-5 dependence · SYNTHETIC · 10,000 runs · horizon T = 10,000 observations · nominal 5%

On correlated runs the estimator fires far too often; on independent scores it holds near nominal. Until a block-aware e-process ships, every certificate states the independence premise on its face. We have not found this number published for any comparable detector — that is exactly why you should demand it.

Start

The record starts when you do.

A monitor set up today is evidence you hold next quarter. Two monitors are free, forever — and the builder sets your first one up with you.