Public incidents require an anytime-valid e-process alarm, the daily e-BH
multiplicity gate, and confound isolation before publication. The certified
design is fixed-panel: one complete, prospectively frozen panel mean is the
only inferential observation.
Where this instrument fails
Under temporally dependent scores — one latent component shared across
5 consecutive observations — the
measured false-alarm rate is
56.38% against a
nominal α = 0.05, over 10,000 seeded runs.
Synthetic Anytime validity requires a
conditional-mean premise, not merely a correct marginal mean. Until a
block-aware e-process or a pre-declared decorrelation rule is available,
certificates attest the computation conditional on that premise;
they do not certify the premise from observed scores.
Details and data ↓
Decision policy
Evidence that stays valid over time
Global α = 0.05
Anytime-valid control
The public default allocates a 5% error budget. Anytime-valid means the
confidence sequence and e-process guarantees remain valid even when
evidence is checked after every probe, rather than only at a fixed,
pre-selected sample size. Degradation is declared by the e-process
alone: an alarm latches when E ≥ 1/α. The confidence-sequence bands
shown on model pages and in certificates are descriptive evidence —
computed over post-activation scores once a baseline freezes — never a
second alarm rule.
δ = 0.03
Practical significance
The default degradation margin is three score points. For baseline μ₀,
the one-sided null is μ ≥ μ₀ − δ, so normal variation inside the margin
is not presented as a meaningful quality decline.
e-BH across streams
Multiplicity gate
The daily e-BH procedure orders stream e-values from largest to
smallest and selects the largest rank k satisfying
k · E(k) / K ≥ 1/α. Only selected streams may publish,
controlling the false discovery rate across the monitored streams.
Offline verification
Certificate verification key
Download a tenant certificate and this public Ed25519 key, then run
actuary verify certificate.json --pubkey actuary-certificate-public-key.txt
without a network connection.
No certificate verification key has been published yet.
Certified design
Fixed-panel validity
This is the design the watchtower runs and certifies. Every figure below
is a simulation of it, loaded from the committed evidence
artifact and produced by the same function the release gate asserts
against — not transcribed by hand. All four are measured over a horizon of
20 complete panels; the design's standing
guarantee is the anytime-valid α = 0.05, and an
empirical false-alarm rate rises toward α as panels accumulate past that
horizon. None of these numbers is a measurement of a live model.
0.252%
False alarms
Null arm — 20 complete panels
Alarming runs
252 of 100,000
95% Wilson upper
0.285%
Guaranteed level
α = 0.05, at any horizon
A rate over 20 complete panels, not a standing
property: it rises toward α as panels accumulate.
100.00%
Interval coverage
Null arm — 20 complete panels
Covering runs
100,000 of 100,000
Null floor
0.70
Panel
5 members, equally weighted
99.58%
Power
Degraded arm — 20 complete panels
Detecting runs
99,580 of 100,000
Alternative
panel mean 0.40 from panel 2 onward
Alarms before the change
0 of 100,000
Power without its alternative and horizon is not a meaningful number, so
both are stated beside it.
Score model: independent Bernoulli per member, 5 members per panel. Both arms run
100,000 replications from a pinned seed, regenerated by
simulate --mode fixed-panel; the server refuses to render this
section if the committed artifact describes any other configuration.
What anytime-validity buys
The cost of an unbudgeted look
The usual way to watch a quality metric is to recompute a test every time
fresh data lands and act when it crosses 0.05. On the degraded lane of the
2026-07-24 study that approach is quicker than ours — it
crosses at block 5, seven panels before the e-process alarms at 12. We draw
it here rather than on the landing page because the honest comparison needs
more than a caption, and because it is the first thing a statistician asks.
The same degraded lane, tested two ways. Each open mark is a look the
design never budgeted for.
Speed is not the claim being contested. A fixed-horizon test spends its 5%
error budget once, at a sample size chosen in advance; recomputing it at
twenty successive blocks spends that budget many times over, so the block-5
crossing carries no error rate anyone can quote. The e-process pays seven
panels of patience and keeps a bound that holds at every look for the life of
the monitor. Both traces are recomputed in the browser from
eprocess.js, which reproduces the signed certificate of study
001 bit for bit; neither is a frequency claim, because one trajectory is not
a rate.
Withdrawn design — historical diagnostics
Null-simulation validity
The cards in this section and the two that follow describe the
withdrawn per-score design, which calibrated a baseline
from observed scores. The certified path above does not calibrate, learn a
baseline, or run a CUSUM. These are retained as diagnostics of that earlier
design, not as evidence for what the watchtower publishes today.
Results below are loaded from the seeded simulation artifact, not copied
into this page. The oracle card fixes μ₀ exactly. The finite-calibration
card is retained as a historical diagnostic of baseline-estimation error;
it does not describe FixedPanelV1, which binds a reviewed null floor before
monitoring and never estimates one from the observed stream.
2.40%
Beta mixture false-alarm rate
Oracle baseline (μ₀ known exactly)
Simulation scale
10,000 null runs × 1,000 observations
Tested α
0.05
α + 2 SE bound
5.44%
Within the stated α-plus-slack validity bound.
3.27%
Beta mixture false-alarm rate
Historical calibration diagnostic (baseline estimated from 3,000 scores)
Simulation scale
10,000 null runs × 1,000 observations
Mean estimated μ₀
0.900
Tested α
0.05
α + 2 SE bound
5.44%
Historical finite-calibration diagnostic only. FixedPanelV1 binds a direct frozen null floor and never learns one from monitored observations.
Historical finite-baseline diagnostic
False-alarm rate by calibration size
Every point uses the same boundary null (μ₀ = 0.90, δ = 0.03,
α = 0.05, T = 1,000). The horizontal reference is the oracle row.
The highlighted n = 3,000 tier records the former calibration design.
It is preserved to show why estimating a floor from a finite monitored
prefix was retired; it is not a current runtime threshold or validity
claim.
Baseline-estimation inflation is material at n = 300 and approaches
the oracle band by the historical n = 3,000 reference tier.
Calibration tier
Boundary false-alarm rate
Runs
Oracle μ₀
2.40%
10,000
n = 300
12.45%
10,000
n = 700
6.89%
10,000
n = 1,500
4.32%
10,000
n = 3,000 (historical reference)
3.27%
10,000
n = 6,000
2.79%
10,000
n = 10,000
2.62%
10,000
n = 20,000
2.52%
10,000
Withdrawn design — historical diagnostics
Power and detection latency
Each point is a harness result for an injected decline beyond the
practical-significance boundary, under the withdrawn per-score design. The
certified path's power is stated in "Fixed-panel validity" above, with its
alternative and horizon.
Post-change detection power by injected effect size.Median observations from changepoint to alarm, conditional on detection.
Effect size
Power
Median latency
Runs
0.020
21.3%
659 observations
10,000
0.050
100.0%
329 observations
10,000
0.100
100.0%
148 observations
10,000
0.200
100.0%
72 observations
10,000
Dependence stress tests
What changes when probes move together
These seeded cells separate the arbitrary cross-stream dependence claim
of e-BH from the per-step conditional-mean floor of each stream's
e-process.
56.38%
Within-stream block stress
10,000 boundary-null runs ×
10,000 observations shared one
latent beta-mixture component for 5
consecutive scores. Positive score correlation makes the next
conditional mean history-dependent, so this is an assumption-sensitivity
result — not an α-control proof.
Gaussian copula ρ
True nulls K₀
Evaluation
Query
Observed score corr.
Empirical FDR
Target + tolerance
Non-null power
Replications
0.0
100
Fixed time
500
-0.000
0.17%
≤ 5.08%
0.0%
10,000
0.0
100
Fixed time
2,000
-0.000
0.13%
≤ 5.07%
0.0%
10,000
0.0
100
Fixed time
10,000
-0.000
0.07%
≤ 5.05%
0.0%
10,000
0.0
100
Each stream's alarm time
τₖ, capped at 10,000
-0.000
0.07%
≤ 5.05%
0.0%
10,000
0.0
90
Fixed time
500
-0.000
0.17%
≤ 4.52%
98.7%
10,000
0.0
90
Fixed time
2,000
-0.000
0.11%
≤ 4.52%
100.0%
10,000
0.0
90
Fixed time
10,000
-0.000
0.06%
≤ 4.51%
100.0%
10,000
0.0
90
Each stream's alarm time
τₖ, capped at 10,000
-0.000
0.06%
≤ 4.51%
90.3%
10,000
0.0
50
Fixed time
500
-0.000
0.10%
≤ 2.51%
99.6%
10,000
0.0
50
Fixed time
2,000
-0.000
0.07%
≤ 2.51%
100.0%
10,000
0.0
50
Fixed time
10,000
-0.000
0.04%
≤ 2.51%
100.0%
10,000
0.0
50
Each stream's alarm time
τₖ, capped at 10,000
-0.000
0.05%
≤ 2.50%
63.0%
10,000
0.5
100
Fixed time
500
0.267
0.11%
≤ 5.07%
0.0%
10,000
0.5
100
Fixed time
2,000
0.267
0.09%
≤ 5.06%
0.0%
10,000
0.5
100
Fixed time
10,000
0.267
0.03%
≤ 5.03%
0.0%
10,000
0.5
100
Each stream's alarm time
τₖ, capped at 10,000
0.267
0.06%
≤ 5.03%
0.0%
10,000
0.5
90
Fixed time
500
0.267
0.16%
≤ 4.53%
98.7%
10,000
0.5
90
Fixed time
2,000
0.267
0.11%
≤ 4.52%
100.0%
10,000
0.5
90
Fixed time
10,000
0.267
0.06%
≤ 4.52%
100.0%
10,000
0.5
90
Each stream's alarm time
τₖ, capped at 10,000
0.267
0.08%
≤ 4.51%
90.2%
10,000
0.5
50
Fixed time
500
0.267
0.10%
≤ 2.51%
99.6%
10,000
0.5
50
Fixed time
2,000
0.267
0.07%
≤ 2.51%
100.0%
10,000
0.5
50
Fixed time
10,000
0.267
0.04%
≤ 2.51%
100.0%
10,000
0.5
50
Each stream's alarm time
τₖ, capped at 10,000
0.267
0.05%
≤ 2.50%
62.0%
10,000
0.9
100
Fixed time
500
0.667
0.07%
≤ 5.05%
0.0%
10,000
0.9
100
Fixed time
2,000
0.667
0.08%
≤ 5.06%
0.0%
10,000
0.9
100
Fixed time
10,000
0.667
0.06%
≤ 5.05%
0.0%
10,000
0.9
100
Each stream's alarm time
τₖ, capped at 10,000
0.667
0.20%
≤ 5.05%
0.0%
10,000
0.9
90
Fixed time
500
0.667
0.16%
≤ 4.55%
98.5%
10,000
0.9
90
Fixed time
2,000
0.667
0.10%
≤ 4.53%
100.0%
10,000
0.9
90
Fixed time
10,000
0.667
0.08%
≤ 4.54%
100.0%
10,000
0.9
90
Each stream's alarm time
τₖ, capped at 10,000
0.667
0.24%
≤ 4.54%
89.8%
10,000
0.9
50
Fixed time
500
0.667
0.11%
≤ 2.52%
99.6%
10,000
0.9
50
Fixed time
2,000
0.667
0.07%
≤ 2.52%
100.0%
10,000
0.9
50
Fixed time
10,000
0.667
0.05%
≤ 2.51%
100.0%
10,000
0.9
50
Each stream's alarm time
τₖ, capped at 10,000
0.667
0.10%
≤ 2.52%
58.8%
10,000
Each row uses bounded Bernoulli score batches generated by a fresh shared
Gaussian factor, then updates the production e-process for every stream.
All current e-values are queried at t = 500, 2,000, and 10,000 and at each
stream's first alarm time τₖ (non-alarms use the cap T = 10,000). Across
10,000 replications, e-BH stays below its α·K₀/K target plus two standard
errors for K₀ = 100, 90, and 50 at ρ = 0.0, 0.5, and 0.9. The observed
score correlation column confirms that ρ changes the joint bounded-score
paths, not merely the seed.
Published probes
What the probes look like
These samples show the task each family poses. They are not the items the
monitors measure. Every epoch measures a declared five-member panel, frozen
before its first observation and committed by the panel_hash in
that epoch's design; those members stay unpublished so they cannot be
optimised against. Only these published samples are rendered here.
Structured output
structured-output-001
Create one library-catalog JSON object from this record: title "Glass Harbor"; author "Mira Sol"; publication year 1998; available true. Return only JSON. It must contain exactly the keys title (string), author (string), publication_year (integer from 1900 through 2100), and available (boolean).
structured-output-041
Convert the transit ticket facts to bare JSON: ticket TK-2026-0410; route R2; zone 1; fare 225 cents; currently valid false. The object must have only ticket_id (pattern TK-2026-four digits), route (string), zone (integer 1 to 4), fare_cents (positive integer), and valid (boolean).
structured-output-071
Return only the enrollment JSON. Student S00841 is enrolled in BIO214 "Field Ecology" for term 2026-fall. Modules in order are core (2 credits) and lab (1 credits); total_credits is their sum. Required shape: student_id string; course object with code and title; term string; modules array of exactly two objects with name and positive integer credits; total_credits positive integer; status one of enrolled, waitlisted, dropped. No extra fields at any level.
structured-output-111
Emit only JSON for experiment EXP-26-101 about soil moisture. Hypothesis text is exactly "treatment improves retention_rate". Cohorts: control size 30, blinded true; treatment size 32, blinded false. Primary metric name retention_rate, unit percent, control_value 40.00, treatment_value 46.25. Quality: randomized true, missing_observations 0. Shape must nest experiment(id/topic/hypothesis), cohorts(control and treatment, each size positive integer and blinded boolean), primary_metric(name/unit/control_value/treatment_value), and quality(randomized/missing_observations). All four root keys are required and no object permits extras.
structured-output-150
Return only strict JSON for batch BAT-2026-05361, product nozzle-J, units 325. Materials in order: code STL-304, mass_kg 38.75, recycled true; code POLY-X, mass_kg 11.95, recycled false. Quality checks: dimensions result pass, sample_size 17; pressure result pass, sample_size 14. Disposition decision release, review_required false, reason_codes exactly []. Shape: batch(id/product/units); materials exactly two code/mass_kg/recycled objects; quality_checks with dimensions and pressure objects, each result pass/fail and positive sample_size; disposition(decision release/hold, review_required boolean, reason_codes array of strings with at most one entry). Forbid unlisted fields at every level.
Extraction
extraction-001
Extract every person name from the note. Return a JSON array of strings only. Note: Amara Voss chaired the review while Julian Pike recorded decisions. The venue was Cedar Hall, and the sponsor was Northwind Labs. A later email mentioned Project Lantern, but no additional person.
extraction-032
Find every complete calendar date in the memo. Return a JSON array of strings exactly as written. The draft was approved on 21 April 2027; publication moved to 2027-05-02. Calls at 09:30 and 16:45 are times, while batch 2036 is only an identifier and not a complete date.
extraction-073
List the telephone numbers in a JSON array exactly as printed. For dispatch call +1 404-555-0171; after hours use (212) 555-0195. Extension 73, case 202555, and postal code 94107 are not complete telephone numbers.
extraction-114
Extract treatments mentioned as actually prescribed and return a JSON array. The care plan starts atorvastatin and schedules cognitive behavioral therapy. The patient asked about aspirin, but it was explicitly not prescribed. Harborlight Clinic and Dr. Wells are not treatments.
extraction-150
Extract all person names, organization names, geographic places, complete calendar dates, email addresses, and uppercase case codes from the dispatch note. Return a JSON array. Note: Jules Park briefed Evergreen Assembly in Ashgrove on 2026-10-19. Replies go to case9@mixed.example, and the filing uses case code MIX-093-K. The room was 404, the call began at 09:15, and the project nickname was Lantern.
Instruction following
instruction-following-001
Write one line about garden sensors. Start exactly with `BRIEF:`, end exactly with `:END`, include `moisture`, use exactly 9 whitespace-separated words, and never use the word `urgent`.
instruction-following-026
Write exactly three nonempty Markdown bullet lines. Include `OPEN` on the first line, `COUNT` on the second, and `CLOSE` on the third. Do not use the word `later`.
instruction-following-056
Return one compact JSON object with exactly the values label=`willow`, count=18, and ready=true. Include no line breaks and do not use the key `comment`.
instruction-following-066
Return a compact JSON array containing exactly the three strings `panda`, `quail`, and `robin` in that order. Keep it on one line and do not include `null`.
instruction-following-139
Write a one-line tagged message that starts with `<INDIA>`, ends with `</INDIA>`, and contains the exact phrase `museum cases inspected`. Its length must be exactly 37 characters. Do not include the word `draft`.
Code reasoning
code-reasoning-001
Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout.
```rust
fn main() {
let values: Vec<i32> = (2..=10).collect();
let kept: Vec<i32> = values
.iter()
.enumerate()
.filter(|(index, value)| (*index as i32 + **value) % 4 != 0)
.map(|(index, value)| *value * (index as i32 % 3 + 1))
.collect();
let alternating: i32 = kept
.iter()
.enumerate()
.map(|(index, value)| if index % 2 == 0 { *value } else { -*value })
.sum();
println!("{kept:?}");
println!("{alternating}");
}
```
code-reasoning-037
Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout.
```rust
use std::collections::BTreeMap;
fn main() {
let updates = [
("amber", 9),
("cobalt", 17),
("amber", -3),
("birch", 6),
("cobalt", 1),
("birch", -1),
];
let mut stock = BTreeMap::new();
for (name, change) in updates {
*stock.entry(name).or_insert(0_i32) += change;
}
stock.retain(|_, count| *count > 0);
for (name, count) in &stock {
println!("{name}={count}");
}
println!("total={}", stock.values().sum::<i32>());
}
```
code-reasoning-075
Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout.
```rust
fn main() {
let left = [6, 9, 12, 8, 11, 7];
let right = [2, 6, 6, 5, 7];
let selected: Vec<i32> = left
.iter()
.zip(right.iter().cycle())
.enumerate()
.filter_map(|(index, (x, y))| {
let value = x * 2 - y + index as i32;
(value % 3 != 0).then_some(value)
})
.collect();
let product = selected.iter().fold(1_i64, |acc, value| acc * i64::from(*value));
println!("{selected:?}");
println!("{product}");
}
```
code-reasoning-113
Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout.
```rust
fn main() {
let value = 252_u16;
let mask = 235_u16;
let mixed = value.rotate_left(3) ^ mask;
let windows: Vec<u16> = (0..4)
.map(|shift| (mixed >> (shift * 2)) & 0b1111)
.collect();
let parity = mixed.count_ones() % 2;
println!("{mixed:016b}");
println!("{windows:?}");
println!("{parity}");
}
```
code-reasoning-150
Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout.
```rust
fn main() {
let mut points = Vec::new();
for row in -3..=3 {
for column in -3..=3 {
let value = row * row + column * column + row * column;
if value % 4 == 0 || row == 0 {
points.push((row, column));
}
}
}
let diagonal = points.iter().filter(|(row, column)| row == column).count();
let balance: i32 = points
.iter()
.enumerate()
.map(|(index, (row, column))| (index as i32 % 3 - 1) * (row - column))
.sum();
println!("count={}", points.len());
println!("diagonal={diagonal}");
println!("balance={balance}");
println!("edges={:?}", (&points[..2], &points[points.len() - 2..]));
}
```
Math
math-001
A supply room has 8 sealed cases with 6 filters per case and 3 loose filters. Every filter costs $4. What is the total cost in dollars? Return only the numeric answer.
math-035
A sensor recorded 70 on 4 trials and 77 on 10 trials. What is the weighted mean across all trials? Round to 3 decimal places and return only the number.
math-073
An arithmetic sequence starts at 9, increases by 5, and contains 14 terms. What is the sum of all terms? Return only the number.
math-109
A bag contains 12 blue tokens and 9 amber tokens. Two tokens are drawn uniformly without replacement. What is the probability, as a percentage, that both are blue? Round to 4 decimal places and return only the number.
math-146
Solve for x: 3.0x + (-1.0) = 17.0. Return only the numeric value of x.