Codex reset track record
Reset Beacon tracks every shared Codex reset, the signals that preceded it, and how the forecast performed against what actually happened.
43 confirmed resets tracked.
8 of 10 explicit reset-day announcements were followed by a reset.
What the record holds
Shared resets only: the extra reset OpenAI hands Codex users generally. Your own five-hour and weekly refills are not counted here.
to , every one carrying the official post it was confirmed from. These are the same events the history page lists day by day.
The typical gap is the middle wait between resets that reached every Codex user: half the waits were shorter and half were longer. Part of the longest wait has no official post on record at all, so it cannot be read as a proven quiet spell.
When a day is named
When OpenAI’s Codex lead names a day for a reset, it has arrived 8 of the last 10 times.
A named day is a public post from OpenAI’s Codex lead that puts a reset on a particular day, counted once per post and scored when the day it named has passed.
What time of day resets land
Each of the 43 confirmed resets counted into the Pacific hour its announcement carries. The hours are Pacific time, OpenAI’s own clock, because an hour-of-day count cannot be moved onto your clock the way a single time can.
The record leans late in the day. 14 of the 43 landed between 15:00–18:00 Pacific time, a stretch that would hold about 7 if resets fell evenly across the day. The 02:00–07:00 Pacific-time stretch holds 0.
This is what happened, not a timetable. A reset arrives at no fixed hour, and the next one can land in any hour above, including one that is empty today.
What has warned before a reset
Two signs beat chance in both halves of the record, measured over the 48 hours after each one. Chance over the same points was 25.4% in the first half and 52.2% in the second, so a rate has to clear both to count. The counts come from the frozen reset record; the methodology page sets out how each one was scored.
| Sign | First half of record | Second half | What it does here |
|---|---|---|---|
| An official post hinting at a reset OpenAI’s Codex lead promises a reset, drops a time-bound hint about one, or answers complaints about drained limits by accepting that a reset is the response. |
4 of 5 times, 80.0% | 3 of 4 times, 75.0% | The only sign strong enough on its own to raise the chance far enough to send an email. |
| The start of an OpenAI status incident OpenAI’s own status page opens an incident that touches Codex. It is the earliest warning that does not depend on anyone posting about resets. |
9 of 25 times, 36.0% | 8 of 12 times, 66.7% | About half of these are followed by nothing, so it nudges the 48-hour chance and never sends an email by itself. |
Fourteen other candidates were scored the same way and dropped: product launches from OpenAI and its rivals, complaint threads on GitHub and Reddit, Fridays, weekends, banked credits, plan changes, and every bucket of time since the last reset. Each one either failed the first half of the record or reversed direction in the second, so none of them moves the number.
Technical scoring
The baseline is the reset rate a forecaster could have known on the day, measured only from windows that had already closed.
Over the last 45 days the forecast has been cautious: it said 34% on average and a reset came 53% of the time. Against the historical rate available at the time, the held-out forecast performed better at both 24 and 48 hours.
The scores below are a replay: today’s method run over the past reset history, forecasting each day with only what was known that day. Every live forecast is frozen the moment it is published and checked when its window closes. Nothing is adjusted afterwards.
Model calibrated-baseline-v1 · 149 replayed days · held-out days to · replay runs through
Calibration by band
A well-calibrated model resets about 40% of the time when it says 40%. We say ‘likely’ only when the chance for that exact window is at least 80%.
The plain average and the daily check
All days since : 32% of 175 measured 48-hour periods contained a reset. That is the plain average a forecast is measured against.
Every published forecast is frozen and checked when its 24- and 48-hour windows close. The replay is the starting point; the live record grows by one day each day and is the record that matters.
How scoring works
Each published reading is frozen with a timestamp. When its window closes, we mark it against the classification record: a reset claimed or confirmed counts as a hit; scheduled-only posts and banked credits do not. Corrections to the classification re-score the affected forecasts and the change is listed in the model change log.
Every scored forecast is downloadable: /api/forecast returns the current reading with its baseline comparison attached.
The full score table
Observed rate is how often a reset really landed in that window; forecast rate is what the model said on average. The training days were used to tune the method; the held-out days were kept aside and never used for tuning, so they are the fairer test. The last column repeats the comparison against the rate those same days turned out to have, which is measured with hindsight nobody had at the time.
| Window | Observed rate | Forecast rate | Brier score | Against the day-of baseline | Against the same days in hindsight | Reading |
|---|---|---|---|---|---|---|
| 24 hours · full window | 18.1% | 12.7% | 0.151 | 0.7% better | 1.6% worse | Brier 0.151 in this slice (0 is perfect, 0.25 is a coin toss). |
| 24 hours · training | 13.5% | 10.3% | 0.119 | 0.1% worse | 2.3% worse | Brier 0.119 in this slice (0 is perfect, 0.25 is a coin toss). |
| 24 hours · held out (last 45 days) | 28.9% | 18.4% | 0.224 | 1.7% better | 8.9% worse | Brier 0.224 in this slice (0 is perfect, 0.25 is a coin toss). |
| 48 hours · full window | 34.9% | 24.4% | 0.234 | 1.6% better | 3.2% worse | Brier 0.234 in this slice (0 is perfect, 0.25 is a coin toss). |
| 48 hours · training | 26.9% | 20.2% | 0.204 | 0.6% better | 3.8% worse | Brier 0.204 in this slice (0 is perfect, 0.25 is a coin toss). |
| 48 hours · held out (last 45 days) | 53.3% | 34.2% | 0.304 | 3.2% better | 22.2% worse | Brier 0.304 in this slice (0 is perfect, 0.25 is a coin toss). |
Brier score: 0 is perfect, 0.25 is what you get by always saying 50%, 1 is perfectly wrong.
What the number means
Every day is scored against what happened. The held-out slice is the last 45 days, never used to fit the model. Against the baseline a forecaster could have known on the day, its Brier score was 0.304 to that baseline's 0.314. Scored instead against the rate those same days turned out to have — a figure nobody had at the time — it was 0.304 to 0.249. Treat a reading as a rough sense of the weather, not a schedule.