Actuary Start free
← Watchtower

Statistical method

Methodology

Public incidents require an anytime-valid e-process alarm, the daily e-BH multiplicity gate, and confound isolation before publication. The certified design is fixed-panel: one complete, prospectively frozen panel mean is the only inferential observation.

Where this instrument fails

Under temporally dependent scores — one latent component shared across 5 consecutive observations — the measured false-alarm rate is 56.38% against a nominal α = 0.05, over 10,000 seeded runs. Synthetic Anytime validity requires a conditional-mean premise, not merely a correct marginal mean. Until a block-aware e-process or a pre-declared decorrelation rule is available, certificates attest the computation conditional on that premise; they do not certify the premise from observed scores. Details and data ↓

Decision policy

Evidence that stays valid over time

Global α = 0.05

Anytime-valid control

The public default allocates a 5% error budget. Anytime-valid means the confidence sequence and e-process guarantees remain valid even when evidence is checked after every probe, rather than only at a fixed, pre-selected sample size. Degradation is declared by the e-process alone: an alarm latches when E ≥ 1/α. The confidence-sequence bands shown on model pages and in certificates are descriptive evidence — computed over post-activation scores once a baseline freezes — never a second alarm rule.

δ = 0.03

Practical significance

The default degradation margin is three score points. For baseline μ₀, the one-sided null is μ ≥ μ₀ − δ, so normal variation inside the margin is not presented as a meaningful quality decline.

e-BH across streams

Multiplicity gate

The daily e-BH procedure orders stream e-values from largest to smallest and selects the largest rank k satisfying k · E(k) / K ≥ 1/α. Only selected streams may publish, controlling the false discovery rate across the monitored streams.

Offline verification

Certificate verification key

Download a tenant certificate and this public Ed25519 key, then run actuary verify certificate.json --pubkey actuary-certificate-public-key.txt without a network connection.

No certificate verification key has been published yet.

Certified design

Fixed-panel validity

This is the design the watchtower runs and certifies. Every figure below is a simulation of it, loaded from the committed evidence artifact and produced by the same function the release gate asserts against — not transcribed by hand. All four are measured over a horizon of 20 complete panels; the design's standing guarantee is the anytime-valid α = 0.05, and an empirical false-alarm rate rises toward α as panels accumulate past that horizon. None of these numbers is a measurement of a live model.

0.252%

False alarms

Null arm — 20 complete panels

Alarming runs
252 of 100,000
95% Wilson upper
0.285%
Guaranteed level
α = 0.05, at any horizon

A rate over 20 complete panels, not a standing property: it rises toward α as panels accumulate.

100.00%

Interval coverage

Null arm — 20 complete panels

Covering runs
100,000 of 100,000
Null floor
0.70
Panel
5 members, equally weighted

99.58%

Power

Degraded arm — 20 complete panels

Detecting runs
99,580 of 100,000
Alternative
panel mean 0.40 from panel 2 onward
Alarms before the change
0 of 100,000

Power without its alternative and horizon is not a meaningful number, so both are stated beside it.

Score model: independent Bernoulli per member, 5 members per panel. Both arms run 100,000 replications from a pinned seed, regenerated by simulate --mode fixed-panel; the server refuses to render this section if the committed artifact describes any other configuration.

What anytime-validity buys

The cost of an unbudgeted look

The usual way to watch a quality metric is to recompute a test every time fresh data lands and act when it crosses 0.05. On the degraded lane of the 2026-07-24 study that approach is quicker than ours — it crosses at block 5, seven panels before the e-process alarms at 12. We draw it here rather than on the landing page because the honest comparison needs more than a caption, and because it is the first thing a statistician asks.

FIXED-HORIZON TEST, RECOMPUTED AT EVERY BLOCKp-VALUE, LOG SCALE1.051e-31e-6CROSSES .05 AT BLOCK 5e-PROCESS ALARMS · PANEL 12PANEL 15101520OPEN MARKS · NO CONTROLLED ERROR RATE AT ANY INTERMEDIATE LOOKFILLED · THE ONE REGISTERED LOOK, n = 20
The same degraded lane, tested two ways. Each open mark is a look the design never budgeted for.

Speed is not the claim being contested. A fixed-horizon test spends its 5% error budget once, at a sample size chosen in advance; recomputing it at twenty successive blocks spends that budget many times over, so the block-5 crossing carries no error rate anyone can quote. The e-process pays seven panels of patience and keeps a bound that holds at every look for the life of the monitor. Both traces are recomputed in the browser from eprocess.js, which reproduces the signed certificate of study 001 bit for bit; neither is a frequency claim, because one trajectory is not a rate.

Withdrawn design — historical diagnostics

Null-simulation validity

The cards in this section and the two that follow describe the withdrawn per-score design, which calibrated a baseline from observed scores. The certified path above does not calibrate, learn a baseline, or run a CUSUM. These are retained as diagnostics of that earlier design, not as evidence for what the watchtower publishes today.

Results below are loaded from the seeded simulation artifact, not copied into this page. The oracle card fixes μ₀ exactly. The finite-calibration card is retained as a historical diagnostic of baseline-estimation error; it does not describe FixedPanelV1, which binds a reviewed null floor before monitoring and never estimates one from the observed stream.

2.40%

Beta mixture false-alarm rate

Oracle baseline (μ₀ known exactly)

Simulation scale
10,000 null runs × 1,000 observations
Tested α
0.05
α + 2 SE bound
5.44%

Within the stated α-plus-slack validity bound.

3.27%

Beta mixture false-alarm rate

Historical calibration diagnostic (baseline estimated from 3,000 scores)

Simulation scale
10,000 null runs × 1,000 observations
Mean estimated μ₀
0.900
Tested α
0.05
α + 2 SE bound
5.44%

Historical finite-calibration diagnostic only. FixedPanelV1 binds a direct frozen null floor and never learns one from monitored observations.

Historical finite-baseline diagnostic

False-alarm rate by calibration size

Every point uses the same boundary null (μ₀ = 0.90, δ = 0.03, α = 0.05, T = 1,000). The horizontal reference is the oracle row. The highlighted n = 3,000 tier records the former calibration design. It is preserved to show why estimating a floor from a finite monitored prefix was retired; it is not a current runtime threshold or validity claim.

False-alarm rate by calibration sizeBoundary-null false-alarm rates for finite baseline estimates, with the oracle-baseline rate shown as a horizontal reference. Boundary false-alarm rate Calibration scores n (log10 scale) 0.0 0.05 0.1 2.5 3.0 3.5 4.0
Baseline-estimation inflation is material at n = 300 and approaches the oracle band by the historical n = 3,000 reference tier.
Calibration tier Boundary false-alarm rate Runs
Oracle μ₀ 2.40% 10,000
n = 300 12.45% 10,000
n = 700 6.89% 10,000
n = 1,500 4.32% 10,000
n = 3,000 (historical reference) 3.27% 10,000
n = 6,000 2.79% 10,000
n = 10,000 2.62% 10,000
n = 20,000 2.52% 10,000

Withdrawn design — historical diagnostics

Power and detection latency

Each point is a harness result for an injected decline beyond the practical-significance boundary, under the withdrawn per-score design. The certified path's power is stated in "Fixed-panel validity" above, with its alternative and horizon.

Detection power by effect sizeArtifact-derived power after the simulated changepoint. Detection power Effect size Δ beyond the degradation boundary 0.0 0.2 0.4 0.6 0.8 1.0 0.05 0.1 0.15 0.2
Post-change detection power by injected effect size.
Detection latency by effect sizeArtifact-derived median observations from changepoint to alarm, conditional on detection. Median detection latency Effect size Δ beyond the degradation boundary 0.0 200.0 400.0 600.0 0.05 0.1 0.15 0.2
Median observations from changepoint to alarm, conditional on detection.
Effect size Power Median latency Runs
0.020 21.3% 659 observations 10,000
0.050 100.0% 329 observations 10,000
0.100 100.0% 148 observations 10,000
0.200 100.0% 72 observations 10,000

Dependence stress tests

What changes when probes move together

These seeded cells separate the arbitrary cross-stream dependence claim of e-BH from the per-step conditional-mean floor of each stream's e-process.

56.38%

Within-stream block stress

10,000 boundary-null runs × 10,000 observations shared one latent beta-mixture component for 5 consecutive scores. Positive score correlation makes the next conditional mean history-dependent, so this is an assumption-sensitivity result — not an α-control proof.

Gaussian copula ρ True nulls K₀ Evaluation Query Observed score corr. Empirical FDR Target + tolerance Non-null power Replications
0.0 100 Fixed time 500 -0.000 0.17% ≤ 5.08% 0.0% 10,000
0.0 100 Fixed time 2,000 -0.000 0.13% ≤ 5.07% 0.0% 10,000
0.0 100 Fixed time 10,000 -0.000 0.07% ≤ 5.05% 0.0% 10,000
0.0 100 Each stream's alarm time τₖ, capped at 10,000 -0.000 0.07% ≤ 5.05% 0.0% 10,000
0.0 90 Fixed time 500 -0.000 0.17% ≤ 4.52% 98.7% 10,000
0.0 90 Fixed time 2,000 -0.000 0.11% ≤ 4.52% 100.0% 10,000
0.0 90 Fixed time 10,000 -0.000 0.06% ≤ 4.51% 100.0% 10,000
0.0 90 Each stream's alarm time τₖ, capped at 10,000 -0.000 0.06% ≤ 4.51% 90.3% 10,000
0.0 50 Fixed time 500 -0.000 0.10% ≤ 2.51% 99.6% 10,000
0.0 50 Fixed time 2,000 -0.000 0.07% ≤ 2.51% 100.0% 10,000
0.0 50 Fixed time 10,000 -0.000 0.04% ≤ 2.51% 100.0% 10,000
0.0 50 Each stream's alarm time τₖ, capped at 10,000 -0.000 0.05% ≤ 2.50% 63.0% 10,000
0.5 100 Fixed time 500 0.267 0.11% ≤ 5.07% 0.0% 10,000
0.5 100 Fixed time 2,000 0.267 0.09% ≤ 5.06% 0.0% 10,000
0.5 100 Fixed time 10,000 0.267 0.03% ≤ 5.03% 0.0% 10,000
0.5 100 Each stream's alarm time τₖ, capped at 10,000 0.267 0.06% ≤ 5.03% 0.0% 10,000
0.5 90 Fixed time 500 0.267 0.16% ≤ 4.53% 98.7% 10,000
0.5 90 Fixed time 2,000 0.267 0.11% ≤ 4.52% 100.0% 10,000
0.5 90 Fixed time 10,000 0.267 0.06% ≤ 4.52% 100.0% 10,000
0.5 90 Each stream's alarm time τₖ, capped at 10,000 0.267 0.08% ≤ 4.51% 90.2% 10,000
0.5 50 Fixed time 500 0.267 0.10% ≤ 2.51% 99.6% 10,000
0.5 50 Fixed time 2,000 0.267 0.07% ≤ 2.51% 100.0% 10,000
0.5 50 Fixed time 10,000 0.267 0.04% ≤ 2.51% 100.0% 10,000
0.5 50 Each stream's alarm time τₖ, capped at 10,000 0.267 0.05% ≤ 2.50% 62.0% 10,000
0.9 100 Fixed time 500 0.667 0.07% ≤ 5.05% 0.0% 10,000
0.9 100 Fixed time 2,000 0.667 0.08% ≤ 5.06% 0.0% 10,000
0.9 100 Fixed time 10,000 0.667 0.06% ≤ 5.05% 0.0% 10,000
0.9 100 Each stream's alarm time τₖ, capped at 10,000 0.667 0.20% ≤ 5.05% 0.0% 10,000
0.9 90 Fixed time 500 0.667 0.16% ≤ 4.55% 98.5% 10,000
0.9 90 Fixed time 2,000 0.667 0.10% ≤ 4.53% 100.0% 10,000
0.9 90 Fixed time 10,000 0.667 0.08% ≤ 4.54% 100.0% 10,000
0.9 90 Each stream's alarm time τₖ, capped at 10,000 0.667 0.24% ≤ 4.54% 89.8% 10,000
0.9 50 Fixed time 500 0.667 0.11% ≤ 2.52% 99.6% 10,000
0.9 50 Fixed time 2,000 0.667 0.07% ≤ 2.52% 100.0% 10,000
0.9 50 Fixed time 10,000 0.667 0.05% ≤ 2.51% 100.0% 10,000
0.9 50 Each stream's alarm time τₖ, capped at 10,000 0.667 0.10% ≤ 2.52% 58.8% 10,000

Each row uses bounded Bernoulli score batches generated by a fresh shared Gaussian factor, then updates the production e-process for every stream. All current e-values are queried at t = 500, 2,000, and 10,000 and at each stream's first alarm time τₖ (non-alarms use the cap T = 10,000). Across 10,000 replications, e-BH stays below its α·K₀/K target plus two standard errors for K₀ = 100, 90, and 50 at ρ = 0.0, 0.5, and 0.9. The observed score correlation column confirms that ρ changes the joint bounded-score paths, not merely the seed.

Published probes

What the probes look like

These samples show the task each family poses. They are not the items the monitors measure. Every epoch measures a declared five-member panel, frozen before its first observation and committed by the panel_hash in that epoch's design; those members stay unpublished so they cannot be optimised against. Only these published samples are rendered here.

Structured output

structured-output-001

Create one library-catalog JSON object from this record: title "Glass Harbor"; author "Mira Sol"; publication year 1998; available true. Return only JSON. It must contain exactly the keys title (string), author (string), publication_year (integer from 1900 through 2100), and available (boolean).

structured-output-041

Convert the transit ticket facts to bare JSON: ticket TK-2026-0410; route R2; zone 1; fare 225 cents; currently valid false. The object must have only ticket_id (pattern TK-2026-four digits), route (string), zone (integer 1 to 4), fare_cents (positive integer), and valid (boolean).

structured-output-071

Return only the enrollment JSON. Student S00841 is enrolled in BIO214 "Field Ecology" for term 2026-fall. Modules in order are core (2 credits) and lab (1 credits); total_credits is their sum. Required shape: student_id string; course object with code and title; term string; modules array of exactly two objects with name and positive integer credits; total_credits positive integer; status one of enrolled, waitlisted, dropped. No extra fields at any level.

structured-output-111

Emit only JSON for experiment EXP-26-101 about soil moisture. Hypothesis text is exactly "treatment improves retention_rate". Cohorts: control size 30, blinded true; treatment size 32, blinded false. Primary metric name retention_rate, unit percent, control_value 40.00, treatment_value 46.25. Quality: randomized true, missing_observations 0. Shape must nest experiment(id/topic/hypothesis), cohorts(control and treatment, each size positive integer and blinded boolean), primary_metric(name/unit/control_value/treatment_value), and quality(randomized/missing_observations). All four root keys are required and no object permits extras.

structured-output-150

Return only strict JSON for batch BAT-2026-05361, product nozzle-J, units 325. Materials in order: code STL-304, mass_kg 38.75, recycled true; code POLY-X, mass_kg 11.95, recycled false. Quality checks: dimensions result pass, sample_size 17; pressure result pass, sample_size 14. Disposition decision release, review_required false, reason_codes exactly []. Shape: batch(id/product/units); materials exactly two code/mass_kg/recycled objects; quality_checks with dimensions and pressure objects, each result pass/fail and positive sample_size; disposition(decision release/hold, review_required boolean, reason_codes array of strings with at most one entry). Forbid unlisted fields at every level.

Extraction

extraction-001

Extract every person name from the note. Return a JSON array of strings only. Note: Amara Voss chaired the review while Julian Pike recorded decisions. The venue was Cedar Hall, and the sponsor was Northwind Labs. A later email mentioned Project Lantern, but no additional person.

extraction-032

Find every complete calendar date in the memo. Return a JSON array of strings exactly as written. The draft was approved on 21 April 2027; publication moved to 2027-05-02. Calls at 09:30 and 16:45 are times, while batch 2036 is only an identifier and not a complete date.

extraction-073

List the telephone numbers in a JSON array exactly as printed. For dispatch call +1 404-555-0171; after hours use (212) 555-0195. Extension 73, case 202555, and postal code 94107 are not complete telephone numbers.

extraction-114

Extract treatments mentioned as actually prescribed and return a JSON array. The care plan starts atorvastatin and schedules cognitive behavioral therapy. The patient asked about aspirin, but it was explicitly not prescribed. Harborlight Clinic and Dr. Wells are not treatments.

extraction-150

Extract all person names, organization names, geographic places, complete calendar dates, email addresses, and uppercase case codes from the dispatch note. Return a JSON array. Note: Jules Park briefed Evergreen Assembly in Ashgrove on 2026-10-19. Replies go to case9@mixed.example, and the filing uses case code MIX-093-K. The room was 404, the call began at 09:15, and the project nickname was Lantern.

Instruction following

instruction-following-001

Write one line about garden sensors. Start exactly with `BRIEF:`, end exactly with `:END`, include `moisture`, use exactly 9 whitespace-separated words, and never use the word `urgent`.

instruction-following-026

Write exactly three nonempty Markdown bullet lines. Include `OPEN` on the first line, `COUNT` on the second, and `CLOSE` on the third. Do not use the word `later`.

instruction-following-056

Return one compact JSON object with exactly the values label=`willow`, count=18, and ready=true. Include no line breaks and do not use the key `comment`.

instruction-following-066

Return a compact JSON array containing exactly the three strings `panda`, `quail`, and `robin` in that order. Keep it on one line and do not include `null`.

instruction-following-139

Write a one-line tagged message that starts with `<INDIA>`, ends with `</INDIA>`, and contains the exact phrase `museum cases inspected`. Its length must be exactly 37 characters. Do not include the word `draft`.

Code reasoning

code-reasoning-001

Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout. ```rust fn main() { let values: Vec<i32> = (2..=10).collect(); let kept: Vec<i32> = values .iter() .enumerate() .filter(|(index, value)| (*index as i32 + **value) % 4 != 0) .map(|(index, value)| *value * (index as i32 % 3 + 1)) .collect(); let alternating: i32 = kept .iter() .enumerate() .map(|(index, value)| if index % 2 == 0 { *value } else { -*value }) .sum(); println!("{kept:?}"); println!("{alternating}"); } ```

code-reasoning-037

Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout. ```rust use std::collections::BTreeMap; fn main() { let updates = [ ("amber", 9), ("cobalt", 17), ("amber", -3), ("birch", 6), ("cobalt", 1), ("birch", -1), ]; let mut stock = BTreeMap::new(); for (name, change) in updates { *stock.entry(name).or_insert(0_i32) += change; } stock.retain(|_, count| *count > 0); for (name, count) in &stock { println!("{name}={count}"); } println!("total={}", stock.values().sum::<i32>()); } ```

code-reasoning-075

Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout. ```rust fn main() { let left = [6, 9, 12, 8, 11, 7]; let right = [2, 6, 6, 5, 7]; let selected: Vec<i32> = left .iter() .zip(right.iter().cycle()) .enumerate() .filter_map(|(index, (x, y))| { let value = x * 2 - y + index as i32; (value % 3 != 0).then_some(value) }) .collect(); let product = selected.iter().fold(1_i64, |acc, value| acc * i64::from(*value)); println!("{selected:?}"); println!("{product}"); } ```

code-reasoning-113

Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout. ```rust fn main() { let value = 252_u16; let mask = 235_u16; let mixed = value.rotate_left(3) ^ mask; let windows: Vec<u16> = (0..4) .map(|shift| (mixed >> (shift * 2)) & 0b1111) .collect(); let parity = mixed.count_ones() % 2; println!("{mixed:016b}"); println!("{windows:?}"); println!("{parity}"); } ```

code-reasoning-150

Predict the exact standard output of the fixed Rust program below. Do not explain your reasoning and do not include code fences; return only the program's stdout. ```rust fn main() { let mut points = Vec::new(); for row in -3..=3 { for column in -3..=3 { let value = row * row + column * column + row * column; if value % 4 == 0 || row == 0 { points.push((row, column)); } } } let diagonal = points.iter().filter(|(row, column)| row == column).count(); let balance: i32 = points .iter() .enumerate() .map(|(index, (row, column))| (index as i32 % 3 - 1) * (row - column)) .sum(); println!("count={}", points.len()); println!("diagonal={diagonal}"); println!("balance={balance}"); println!("edges={:?}", (&points[..2], &points[points.len() - 2..])); } ```

Math

math-001

A supply room has 8 sealed cases with 6 filters per case and 3 loose filters. Every filter costs $4. What is the total cost in dollars? Return only the numeric answer.

math-035

A sensor recorded 70 on 4 trials and 77 on 10 trials. What is the weighted mean across all trials? Round to 3 decimal places and return only the number.

math-073

An arithmetic sequence starts at 9, increases by 5, and contains 14 terms. What is the sum of all terms? Return only the number.

math-109

A bag contains 12 blue tokens and 9 amber tokens. Two tokens are drawn uniformly without replacement. What is the probability, as a percentage, that both are blue? Round to 4 decimal places and return only the number.

math-146

Solve for x: 3.0x + (-1.0) = 17.0. Return only the numeric value of x.