Changing a baseline after the fact is moving the yardstick, and a prediction whose yardstick moves is not a prediction. The rule adopted instead is procedural: when the meeting resolves, both benchmarks are reported together with their as-of dates, and the baselines are refreshed before anything new is recorded. The direction of the error is stated rather than corrected - if the true market price on 1 August sat closer to a hike than the 3 July monitor did, the model was further from the market than this record says, not nearer.
ledger.py score on the frozen record, not on a
re-reading of it.The squared error summed across the whole action space, not just the call that turned out right: a confident right answer and a hedged right answer are not the same thing, and a distribution that hedged everything should not score like one that committed. Hold 75 / hike 25 scores 0.125 if it holds and 1.125 if it hikes. Lower is better. Below 0.25 beats always saying fifty-fifty, which is the only bar this machine has to clear to have said anything at all.
The watch list is three names, written down in advance. Precision is how many of those three actually dissented; recall is how many of the actual dissenters were on the list. They are reported separately on purpose: naming everybody would give perfect recall and worthless precision, and naming nobody would do the reverse. This is the part of the output a rates market does not produce, so it is the part worth being measured on.
The engine does not grade its own comparison. The market benchmark is entered with its date, the outcome is entered with its date, and whether one beat the other is written down as a sentence rather than computed into a flattering number.
| Meeting | Predicted | Actual | Brier | Dissent P/R | Beat market |
|---|---|---|---|---|---|
| No scored predictions yet. The first one lands on 16 September 2026. | |||||
It retires exactly one sentence - “never scored” becomes “scored once”. One meeting is not a track record either way, and a right answer for a wrong reason will be published as such: the two checks below already name the mechanism that would make a correct hold suspicious.
The Chair persona was distilled on 1 August from documents that existed then; the keynote did not exist until 28 August, so the model could not have read it. 5 substantive matches, of which the strongest is arithmetic rather than phrasing: the model counted the overshoot to “month 65 or 66” and he said “65 months”. Against that, one pillar failed outright. The model had the Chairman holding because the market had already delivered the tightening; he said in the keynote that he would be hard pressed to describe financial conditions as restrictive at all. That is not a shade of difference - it is the mechanism the hold rested on, gone.
Three days later the first reading was overturned - not the observation, the cause. The model did not invent the pillar. It is the Chairman's own mechanism, from his 29 July press conference, and it was in the corpus the card was built from. Two errors, both of them the kind that survives a spell-check: A hawkish argument was used as a dovish one - the sentence he actually said was vouching for transmission, not explaining why he could wait. And A change was promoted into a level - on 29 July he was describing the intermeeting change in conditions, while on 17 June and 28 August he was describing the level. The model read one as the other.
Nothing. The record for September is already closed, and editing it because new information arrived would mean it was never a prediction. What the re-diagnosis changes is the reading, and the reading tilts further toward a hike than the first one did. Both documents are reported alongside the result when the meeting resolves; if the outcome is a hike, this is the ready-made explanation of why the model was wrong.