Reset Beacon
MODEL ACCURACY

Accuracy record

Experimental — no skill yet

Every forecast is stored the moment it is published and scored once the window closes. Nothing here is adjusted afterwards, except when a source classification is corrected — that re-scores the affected forecasts and is listed in the change log.

Model empirical-hazard-v4 · 149 scored days · scored through August 25, 2026

Scores by window

The fixed rate is the same number every day — the measured base rate for that window. Skill is the percentage by which the model beats it; a negative number means the fixed rate would have served you better.

WindowObserved rateForecast rateBrier scoreSkill vs. fixed rateReading
24 hours · full window 18.1% 18.5% 0.147 +1.2% Better than the fixed rate in this slice.
24 hours · training 13.5% 15.4% 0.118 −1.1% The fixed rate was better in this slice.
24 hours · held out (last 45 days) 28.9% 25.6% 0.213 −3.8% The fixed rate was better in this slice.
48 hours · full window 34.9% 33.1% 0.224 +1.2% Better than the fixed rate in this slice.
48 hours · training 26.9% 28.3% 0.200 −1.9% The fixed rate was better in this slice.
48 hours · held out (last 45 days) 53.3% 44.1% 0.280 −12.4% The fixed rate was better in this slice.

Brier score: 0 is perfect, 0.25 is what you get by always saying 50%, 1 is perfectly wrong.

What the number means

0.224
48-hour Brier score

Better than a coin flip. Slightly better than repeating 32% every day over the full record, and worse than it on the held-out slice. Treat a reading as a rough sense of the weather, not a schedule.

−12.4%
Skill against the fixed rate

Negative, so the model has earned no trust yet. It stays labelled experimental until it beats the fixed rate in one backtest, on both the training and the held-out window, for both horizons.

32%
Measured 48-hour base rate

32% of 175 measured 48-hour windows contained a broad reset. The experimental model has not beaten that simple rate in the held-out record.

Calibration by band

A well-calibrated model resets about 40% of the time when it says 40%.

10–20% · 20 days
15% actually reset
20–30% · 44 days
32% actually reset
30–40% · 46 days
39% actually reset
40–50% · 26 days
46% actually reset
50–60% · 13 days
38% actually reset
What we saidWhat happened

How scoring works

Each published reading is frozen with a timestamp. When its window closes, we mark it against the classification record: a broad reset claimed or confirmed counts as a hit; scheduled-only posts and banked credits do not. Corrections to the classification re-score the affected forecasts and the change is listed in the model change log.

Read the methodologySee the sourcesScored forecasts via API