Methodology

Evidence first. Forecast second.

The system records what a source actually establishes, preserves uncertainty, and withholds probability claims until they beat a simple baseline in repeated evaluations.

The core rule

Timing and proof are different questions.

A statement that a reset is planned can support a forecast. It cannot prove the action happened. An official completion post can establish that the action was triggered. It cannot prove the change reached every account. Reset Beacon tracks these as separate states.

When a required source check expires, the last estimate is labeled paused. A failed source never counts as evidence that nothing is happening.

Forecast model

What the 24 and 48 hour percentages calculate.

The target is at least one automatically applied broad Codex weekly-quota reset in a rolling horizon.

Eligible history

The baseline uses curated official broad-reset completion claims. Scope-unverified acknowledgements, banked credits, personal five-hour windows, natural weekly resets, policy changes, temporary boosts, narrow plan events and unconfirmed schedules are excluded.

Baseline

Completed events are weighted with a 60-day half-life. The weighted daily event rate is converted to each horizon with p(H) = 1 - exp(-rate × H). The model does not raise the chance merely because a reset feels due.

After a reset

A documented cooldown reduces the chance of another immediate reset. The multiplier is 0.12 for the first 12 hours, 0.20 through 24 hours, 0.45 through 48 hours, 0.70 through 96 hours, 0.90 through seven days, then 1.00.

Direct official statements

An explicit intention sets documented floors of 55% for 24 hours and 75% for 48 hours. An official schedule inside the horizon sets floors of 92% and 96%. Scheduled delivery can slip, so no future event is shown as 100%.

Audit trail

Each calculation stores the model version and hash, source health, exact basis-point contributions, previous and new displayed values, expiry, reason and source. The rows must add exactly to the published estimate or the snapshot is rejected.

State model

What each label establishes.

The public label advances only when evidence meets the rule for that state. Corrections can move an event backward.

No current signal
No fresh evidence supports a special reset event.
Unverified report
A claim exists, but its origin, context or independence is not strong enough for scheduling.
Official intent
An authorized source states that a reset is planned, without a usable time.
Officially scheduled
An official source gives a date, time or bounded window. The original wording and timezone are retained.
Action claimed
An official source says the reset action completed. Account propagation may still be pending.
Rollout observed
Fresh, non-duplicate paid-account reports show before-and-after changes across more than one plan and region. Sample limits remain visible.
Account confirmed
A user’s fresh account reading shows a material allowance change outside its natural reset and without consuming a banked credit.
Window passed, unconfirmed
The scheduled window ended without enough action or rollout evidence.
Corrected or cancelled
A source changed the timing, withdrew the plan, or later evidence invalidated the event.

Time handling

One instant, three useful views.

Every machine-readable time is stored as an exact UTC instant. The page then renders it in the visitor’s device timezone with Intl.DateTimeFormat. The source timezone remains directly below it, and UTC is available as an audit fallback.

No Chicago defaultA visitor in California sees Pacific local time. A visitor in Chicago sees Central local time. If a source writes “PST” during daylight time, the literal offset and likely local-time interpretation remain distinguishable until clarified.

Evidence weight

Source strength is not a vote count.

Ten reposts of one claim remain one underlying source. We cluster duplicates before they can affect the state.

Evidence classTypical useCannot prove
Official explicit timeScheduling a reset windowCompletion or account propagation
Official clarificationResolving date, hour or wordingThat the action happened
Official completion statementAdvancing to action claimedThat every account updated
Fresh account readingPersonal account confirmationUniversal rollout by itself
Independent paid-user reportsMeasuring rollout breadthOfficial scheduling
Status incidents and historical patternsContext and prior probabilityA special reset commitment

Accuracy

Scores must earn their place.

The current number is labeled experimental. It is not advertised as an accuracy rate or an OpenAI commitment. Frozen walk-forward evaluations will publish Brier score, log loss, calibration, a constant-rate baseline, false alerts, missed events, sample size and model version.

The experimental label remains until the model beats its baseline in two consecutive frozen backtests. Alert thresholds are not activated merely because an unvalidated model crosses a number.

Corrections

Source edits, timezone ambiguity and missed windows remain in the event record. Corrections are sent to subscribers who received the affected alert.