Actuary

How it works

A day in the life of one monitor.

First we agree three things and write them down: which test cases to use, the lowest score you are willing to live with, and how often you would tolerate a false alarm. After that, every day is the same — run the tests, take one score, add it to the evidence. That really is the whole loop.

Start free See one check run

Agreed and locked before any data

5 TEST CASES · LOWEST SCORE YOU’LL ACCEPT 0.70 · AT MOST 1 FALSE ALARM IN 20

  1. Run the test cases all five of them, sent and marked
  2. Take one score their average, 0.88 — nothing else counts
  3. Update the evidence now at 0.79, and an alarm needs 20

↺ tomorrow, with exactly the same rules

Day 7 of 7 · no alarm, and a quiet week is worth recording too

AGREED ONCE, THEN MEASURED THE SAME WAY EVERY DAY

Day one

We build the first monitor with you.

STEP 1

Name the task

The thing your feature must keep doing — extract the fields, keep the format, get the number right. “Extract the invoice fields accurately” is a task; “monitor GPT-5” is not. A task is narrow enough when one score means the same thing from one check to the next.

STEP 2

Write the rules down

The test cases, the lowest score you’ll accept, the false-alarm budget — all settled before the first result arrives. Set the minimum from what the product needs, not what the model scores today: too low and real damage sails under it, too high and you live at the edge of alarm. Change any rule later and that starts a new chapter; the old one stays visible beside it.

STEP 3

Watch it fill up

One page per monitor, showing every check, how sure we currently are, and every change anyone has made. Nothing is ever added after the fact.

Two monitors are free, forever. We do the setup with you rather than handing you a form and wishing you luck.

What runs

If we do the testing, here is what we test.

Five kinds of task: structured output, pulling fields out of text, following instructions, reasoning about code, and maths. Each one has a pool of 150 questions whose answers we already know, and every reply is marked by ordinary code that gives it a score between 0 and 1. We never ask one model to judge another.

A monitor draws five questions from that pool and fixes them in place before the first result arrives. We keep the exact five to ourselves, so that nobody — including us — can quietly tune against them. The two examples below come from the set we publish on purpose.

math-001 · PUBLISHED SAMPLE

A supply room has 8 sealed cases with 6 filters per case and 3 loose filters. Every filter costs $4. What is the total cost in dollars? Return only the numeric answer.

instruction-following-001 · PUBLISHED SAMPLE

Write one line about garden sensors. Start exactly with BRIEF:, end exactly with :END, include moisture, use exactly 9 whitespace-separated words, and never use the word urgent.

The check

One check, from start to finish.

We set the money aside before making a single call to your provider; if there isn’t any, we make no calls at all rather than half a check. Every test case comes back with a score, and one that fails outright scores zero — we never quietly drop it. Only a complete set of five moves the evidence along. Look at the result once a year or twenty times a day; the promise already covers every look.

MONEY SET ASIDE FIRST · NO BUDGET MEANS NO CALLS AT ALL 1.0 1.0 0.8 1.0 0 THE SAME FIVE TEST CASES, EVERY TIME FAILED OUTRIGHT → SCORES 0, NEVER DROPPED AVERAGE · 0.76 THE ONLY NUMBER THAT COUNTS ALARM LEVEL · 20 EVIDENCE BUILDS UP, ONE CHECK AT A TIME
ONE CHECK, START TO FINISH · A TEST CASE THAT FAILS OUTRIGHT SCORES 0 AND IS NEVER DROPPED

On the public record we run one check per model per day. If we host the checks for you, they run as often as your setup says. What the promise assumes, and where it stops holding: read the boundary.

Your scores

Already scoring your feature? Just send us the number.

Tell us which of our built-in checkers to use and how to configure it — settings only, never code we run for you — then send one score per check. An eval harness, a human review queue, a metric you already trust: if it comes out as a number between 0 and 1, we can watch it.

PUT /v1/suites/checkout · application/yaml

checker: math
parameters:
  absolute_tolerance: 0.01
  relative_tolerance: 0.001

POST /v1/scores

{"suite":"support-agent","suite_version":3,
 "model":"gpt-5-mini","score":0.86,
 "ts":"2026-07-29T09:14:07Z"}

Scores you send us

You get the full history, the API and a page on the record. We label these monitors "descriptive", because the rules were not agreed in advance — which means we won’t put a certificate behind them, and we’d rather say so than blur the line.

WE ONLY STORE THE NUMBER AND ITS LABELS — NEVER YOUR PROMPTS OR OUTPUTS

Designs we freeze together

Sit down with us, agree the test cases and the minimum score, and have it reviewed. From then on the same machinery carries the published false-alarm promise and every alarm arrives with a signed certificate.

See what a certificate carries

The alarm

An alarm arrives with its homework done.

Once the evidence passes the alarm level it stays there — it cannot un-fire itself. You get one message in Slack and email naming the monitor, how strong the evidence is, and when we think the trouble started, with a link to the incident page. That page keeps the whole story, and on designs we froze with you it also carries a signed certificate anyone can check offline. Quiet periods are published just as carefully: "no alarm" is a result, not an absence.

THE ALERT Slack and email · which monitor · how strong · when it started OPENS THE INCIDENT the whole story, kept on the public record ON DESIGNS WE FROZE WITH YOU, ALSO THE CERTIFICATE signed, and anyone can recheck it offline EVERYTHING YOU NEED, NOT JUST A PING

STUDY 001 · 24 JULY 2026 · ALARM AT CHECK 12 · WE HAD PROMISED WITHIN 19 · ONE REAL RUN, NOT A SUCCESS RATE

This is the only detection on the record so far, and we caused it deliberately to prove the whole path works. Read study 001

The dials

You choose the settings. Then we all live with them.

Setting What you can choose
How answers are marked one of five built-in checkers, configured with settings rather than code · every score records which version marked it
Checker settings how close a number has to be · whether case matters · which constraints to enforce · how code is normalised · a JSON Schema you supply
The test cases five drawn from the pool, fixed in place by hash before the first result
Minimum score and false-alarm budget set per monitor, before any data · by default we allow at most one false alarm in twenty
How often it runs whatever schedule you write into the design
Where alerts go Slack and email · send yourself a test alert whenever you like
Stopping and starting pause or resume a monitor, and change its settings, through the API

On the public record we watch open-weight models such as Llama 3.3 70B and Qwen 3 30B alongside the frontier ones, and we always name the exact route: model, provider, and which version of the weights. The Watchtower shows who is being watched right now, and you can ask us to add a model you care about. Request a model →

Your stack

Where it sits in your stack.

Tracing shows what happened inside one run. Evaluation defines the checks and produces the scores. Actuary sits after both: it watches the scores over time and decides when the pattern is strong enough to call a production regression. Use it beside LangSmith, Phoenix, Braintrust, OpenTelemetry or your own eval pipeline. Anything that produces one trustworthy number per check will do. More on the split: monitoring vs observability.

Start

Two monitors free, forever.

Send us one score today and you have started. Or just describe the task, and we’ll come back with a monitor already set up — test cases, minimum score and false-alarm budget, all written down where you can see them.

Start free Apply for a design partnership