Accuracy record
Every forecast is stored the moment it is published and scored once the window closes. Nothing here is adjusted afterwards, except when a source classification is corrected — that re-scores the affected forecasts and is listed in the change log.
Model empirical-hazard-v4 · 149 scored days · scored through August 25, 2026
Scores by window
The fixed rate is the same number every day — the measured base rate for that window. Skill is the percentage by which the model beats it; a negative number means the fixed rate would have served you better.
| Window | Observed rate | Forecast rate | Brier score | Skill vs. fixed rate | Reading |
|---|---|---|---|---|---|
| 24 hours · full window | 18.1% | 18.5% | 0.147 | +1.2% | Better than the fixed rate in this slice. |
| 24 hours · training | 13.5% | 15.4% | 0.118 | −1.1% | The fixed rate was better in this slice. |
| 24 hours · held out (last 45 days) | 28.9% | 25.6% | 0.213 | −3.8% | The fixed rate was better in this slice. |
| 48 hours · full window | 34.9% | 33.1% | 0.224 | +1.2% | Better than the fixed rate in this slice. |
| 48 hours · training | 26.9% | 28.3% | 0.200 | −1.9% | The fixed rate was better in this slice. |
| 48 hours · held out (last 45 days) | 53.3% | 44.1% | 0.280 | −12.4% | The fixed rate was better in this slice. |
Brier score: 0 is perfect, 0.25 is what you get by always saying 50%, 1 is perfectly wrong.
What the number means
Better than a coin flip. Slightly better than repeating 32% every day over the full record, and worse than it on the held-out slice. Treat a reading as a rough sense of the weather, not a schedule.
Negative, so the model has earned no trust yet. It stays labelled experimental until it beats the fixed rate in one backtest, on both the training and the held-out window, for both horizons.
32% of 175 measured 48-hour windows contained a broad reset. The experimental model has not beaten that simple rate in the held-out record.
Calibration by band
A well-calibrated model resets about 40% of the time when it says 40%.
How scoring works
Each published reading is frozen with a timestamp. When its window closes, we mark it against the classification record: a broad reset claimed or confirmed counts as a hit; scheduled-only posts and banked credits do not. Corrections to the classification re-score the affected forecasts and the change is listed in the model change log.