math-001 · PUBLISHED SAMPLE
A supply room has 8 sealed cases with 6 filters per case and 3 loose filters. Every filter costs $4. What is the total cost in dollars? Return only the numeric answer.
How it works
First we agree three things and write them down: which test cases to use, the lowest score you are willing to live with, and how often you would tolerate a false alarm. After that, every day is the same — run the tests, take one score, add it to the evidence. That really is the whole loop.
Agreed and locked before any data
5 TEST CASES · LOWEST SCORE YOU’LL ACCEPT 0.70 · AT MOST 1 FALSE ALARM IN 20
↺ tomorrow, with exactly the same rules
Day 7 of 7 · no alarm, and a quiet week is worth recording too
Day one
STEP 1
The thing your feature must keep doing — extract the fields, keep the format, get the number right. “Extract the invoice fields accurately” is a task; “monitor GPT-5” is not. A task is narrow enough when one score means the same thing from one check to the next.
STEP 2
The test cases, the lowest score you’ll accept, the false-alarm budget — all settled before the first result arrives. Set the minimum from what the product needs, not what the model scores today: too low and real damage sails under it, too high and you live at the edge of alarm. Change any rule later and that starts a new chapter; the old one stays visible beside it.
STEP 3
One page per monitor, showing every check, how sure we currently are, and every change anyone has made. Nothing is ever added after the fact.
Two monitors are free, forever. We do the setup with you rather than handing you a form and wishing you luck.
What runs
Five kinds of task: structured output, pulling fields out of text, following instructions, reasoning about code, and maths. Each one has a pool of 150 questions whose answers we already know, and every reply is marked by ordinary code that gives it a score between 0 and 1. We never ask one model to judge another.
A monitor draws five questions from that pool and fixes them in place before the first result arrives. We keep the exact five to ourselves, so that nobody — including us — can quietly tune against them. The two examples below come from the set we publish on purpose.
math-001 · PUBLISHED SAMPLE
A supply room has 8 sealed cases with 6 filters per case and 3 loose filters. Every filter costs $4. What is the total cost in dollars? Return only the numeric answer.
instruction-following-001 · PUBLISHED SAMPLE
Write one line about garden sensors. Start exactly with BRIEF:,
end exactly with :END, include moisture, use exactly
9 whitespace-separated words, and never use the word urgent.
The check
We set the money aside before making a single call to your provider; if there isn’t any, we make no calls at all rather than half a check. Every test case comes back with a score, and one that fails outright scores zero — we never quietly drop it. Only a complete set of five moves the evidence along. Look at the result once a year or twenty times a day; the promise already covers every look.
On the public record we run one check per model per day. If we host the checks for you, they run as often as your setup says. What the promise assumes, and where it stops holding: read the boundary.
Your scores
Tell us which of our built-in checkers to use and how to configure it — settings only, never code we run for you — then send one score per check. An eval harness, a human review queue, a metric you already trust: if it comes out as a number between 0 and 1, we can watch it.
PUT /v1/suites/checkout · application/yaml
checker: math parameters: absolute_tolerance: 0.01 relative_tolerance: 0.001
POST /v1/scores
{"suite":"support-agent","suite_version":3,
"model":"gpt-5-mini","score":0.86,
"ts":"2026-07-29T09:14:07Z"}
You get the full history, the API and a page on the record. We label these monitors "descriptive", because the rules were not agreed in advance — which means we won’t put a certificate behind them, and we’d rather say so than blur the line.
WE ONLY STORE THE NUMBER AND ITS LABELS — NEVER YOUR PROMPTS OR OUTPUTS
Sit down with us, agree the test cases and the minimum score, and have it reviewed. From then on the same machinery carries the published false-alarm promise and every alarm arrives with a signed certificate.
The alarm
Once the evidence passes the alarm level it stays there — it cannot un-fire itself. You get one message in Slack and email naming the monitor, how strong the evidence is, and when we think the trouble started, with a link to the incident page. That page keeps the whole story, and on designs we froze with you it also carries a signed certificate anyone can check offline. Quiet periods are published just as carefully: "no alarm" is a result, not an absence.
STUDY 001 · 24 JULY 2026 · ALARM AT CHECK 12 · WE HAD PROMISED WITHIN 19 · ONE REAL RUN, NOT A SUCCESS RATE
This is the only detection on the record so far, and we caused it deliberately to prove the whole path works. Read study 001
The dials
| Setting | What you can choose |
|---|---|
| How answers are marked | one of five built-in checkers, configured with settings rather than code · every score records which version marked it |
| Checker settings | how close a number has to be · whether case matters · which constraints to enforce · how code is normalised · a JSON Schema you supply |
| The test cases | five drawn from the pool, fixed in place by hash before the first result |
| Minimum score and false-alarm budget | set per monitor, before any data · by default we allow at most one false alarm in twenty |
| How often it runs | whatever schedule you write into the design |
| Where alerts go | Slack and email · send yourself a test alert whenever you like |
| Stopping and starting | pause or resume a monitor, and change its settings, through the API |
On the public record we watch open-weight models such as Llama 3.3 70B and Qwen 3 30B alongside the frontier ones, and we always name the exact route: model, provider, and which version of the weights. The Watchtower shows who is being watched right now, and you can ask us to add a model you care about. Request a model →
Your stack
Tracing shows what happened inside one run. Evaluation defines the checks and produces the scores. Actuary sits after both: it watches the scores over time and decides when the pattern is strong enough to call a production regression. Use it beside LangSmith, Phoenix, Braintrust, OpenTelemetry or your own eval pipeline. Anything that produces one trustworthy number per check will do. More on the split: monitoring vs observability.
Start
Send us one score today and you have started. Or just describe the task, and we’ll come back with a monitor already set up — test cases, minimum score and false-alarm budget, all written down where you can see them.