1 · TARGET
What the model predicts
The target is the first public announcement of a new broad, free, automatic Codex reset inside the next 24 or 48 hours. Both horizons are shown together. The 48-hour probability must include the 24-hour opportunity, so it cannot be smaller when both values come from the same snapshot.
The target is not a visitor’s personal weekly timer. It is also not a purchased reset, a banked grant that can be redeemed later, an ordinary quota refresh, a temporary rate increase or the resolution of an outage. A separate planned notice can be displayed prominently, but it does not prove that every account received anything.
2 · EVIDENCE
What evidence is watched
The strongest evidence is original, attributable information: official OpenAI guidance, public announcements from responsible product leaders and a manually reviewed event history. Relevant status incidents and public product signals can provide context, but an outage by itself is not labeled as a reset.
Every candidate event needs a canonical identity so reposts do not multiply the count. The ledger keeps the source URL, publication and observation times, scope, date precision, verification state and whether the event is eligible for the global model. Source coverage is itself evidence; a missing row during an unaudited period is unknown, not a confirmed non-event.
3 · CALCULATION
How probability is calculated
The reference model begins with a conservative event-rate baseline. In plain language, it asks how often eligible resets occurred during audited time and uses that rate to estimate the chance of at least one event during the chosen horizon. A Bayesian prior prevents a tiny sample from creating an extreme answer.
The cadence feature then compares today’s elapsed wait with historical gaps that were still ongoing at the same point. Gaps that had already ended are not treated as survivors. This can adjust the baseline, but only after it has been tested on later, unseen periods. Both horizons use one immutable source snapshot, and a stale or failed source check suppresses the published numbers.
4 · EXCLUSIONS
What does not automatically increase the score
Random tweets, repeated reposts, unsourced screenshots and community rumors do not become ground truth. A single service incident, slow response, ordinary usage-window reset or anecdote from one account does not automatically raise the global-reset probability. Targeted announcements can appear in the timeline while remaining excluded from the broad model.
Human review is not a cosmetic label. Conflicting times require resolution, retrospective corrections cannot leak into older forecasts, and policy changes may create a new modeling era rather than being mixed blindly with older behavior.
5 · LIMITS
Accuracy and limitations
OpenAI can change product policy at any time, and public reset history may remain sparse. A neat pattern in a few gaps can be coincidence. The model therefore needs walk-forward evaluation against simple baselines, calibration checks and enough audited exposure before publication. Event count matters more than the number of hourly rows generated from the same few events.
Even a well-calibrated probability can be wrong on a specific day. “30%” would mean that comparable predictions should occur roughly three times in ten—not that a reset is thirty percent complete. If the model has not earned a number, the honest display is an em dash with a reason and the last successful source-check time.