MLB card — 2026-10-07

Wednesday 2026-10-07 slate
POSTSEASON (AL Division Series, NL Division Series): the model has never seen a postseason game, same methodology. 5 notice(s) below.

PRICES ON THIS CARD ARE NOT DECISION-TIME PRICES. The snapshot behind them was collected at 20:00Z, which is 4h 30m after the 15:30Z decision moment (tolerance 15m). Every price, break-even and edge below is as of 20:00Z, not 15:30Z, and the model saw 4h 30m of information a card built on time could not have had. Treat this slate as unpriced for any claim about beating the market. NOTHING ON THIS CARD WAS RECORDED TO THE LIVE LEDGER — past 90 minutes the picks are not ours to claim we committed to. Read them and bet them if you want; they will not appear in the season record either way.

1 of 4 games had already started when this card was built and were NOT published to the live record — a pick written down after first pitch is not evidence that we committed in advance. They are shown here for reading only.

THE PITCHER DATA IS 10 DAYS OUT OF DATE (last game recorded 2026-09-27). Every starter looks 10 days more rested than he is, which inflates 'days since he last pitched' and can wrongly mark a starter as returning from injury. Twenty-four of the model's fifty-four inputs come from that file, so tonight's probabilities are affected, not just the HIGH/MEDIUM/LOW labels. Run the refresh-gamelogs workflow before trusting this card.

This cost this slate: its decision snapshot landed 4h 30m after the decision time, so the external trigger did not cover the gap either

GITHUB IS DROPPING SCHEDULED RUNS — delivered today 1 of 8 due so far, yesterday 7 of 14. Worst: late-collection 3/11.
GITHUB IS DROPPING SCHEDULED RUNS: today only 1 of the 8 scheduled runs due so far were DELIVERED, yesterday only 7 of 14 were delivered. The biggest shortfall is late-collection: 3 delivered of 11 due. The card may still look correct — on 2026-08-26 it was, and the pipeline failed the next day. Nothing needs fixing in this repo; this is GitHub delivering less than the schedule asks for. THIS IS A COUNT OF RUNS, NOT OF DATA: a poll that stands down exits zero and is counted here as delivered, so these figures can look survivable while a ledger receives nothing at all. Whether anything was actually written is the separate line from src/collection_health.py.

4 games · times in Phoenix · generated 2026-10-07 20:05 UTC
Side is the only validated column. Total* and NRFI* are not — see the notes at the bottom.
Probabilities: calibrated on 2,155 games from the prior year.

Season to date — three markets, never pooled
Every published pickGood data (Aug 18–26, Sep 1 →)Known-bad inputs (Aug 10–17, Aug 27–28)
No-model base rate — what backing the same side of every game would hit
MONEYLINE 53.2% — backing the home team every game
TOTALS 51.0% — backing the under every game
NRFI 50.6% — picking NRFI every game

ahead   behind   against the no-model base rate for that market (above). Units against zero. The triangles carry the same meaning as the colours.

Every published pick — LIVE: known-bad inputs Aug 10-17, Aug 27-28, good data Aug 18-26, Sep 1 onward
▸ Show why 2026-08-10 to 09-03 used a frozen game history

2026-08-10 to 2026-09-03: built with a game history frozen at 2026-08-05, so 14 of the model's 54 features described a 10- and 30-game window that had rolled over, and 6 more carried a stale roster-move count. Not tagged; these picks are counted in every group above. Re-run with the corrected history, 23.9% of picks flipped side.

Every pick published, both groups added together, regardless of data quality. The good-data and known-bad groups are one click away.

MONEYLINE
 nrecordhit %95% CIunitsb/e
all592344-24858.1% 54.1% – 62.0%+18.7 56.5%
HIGH3826-1268.4% 52.5% – 80.9%+1.7 64.9%
MEDIUM206130-7663.1% 56.3% – 69.4%+12.8 59.0%
LOW348188-16054.0% 48.8% – 59.2%+4.1 54.1%
TOTAL* TESTED & FAILED — no edge over 8,549 past games.
 nrecordhit %95% CIunitsb/e
all531265-26649.9% 45.7% – 54.1%-27.4 52.6%
strong9150-4154.9% 44.7% – 64.8%+2.8 53.1%
mid10452-5250.0% 40.6% – 59.4%-5.2 52.5%
weak336163-17348.5% 43.2% – 53.8%-25.0 52.4%
NRFI* UNGRADED — no price history exists to check it against.
 nrecordhit %95% CI
all592294-29849.7% 45.6% – 53.7%
strong14167-7447.5% 39.5% – 55.7%
mid225112-11349.8% 43.3% – 56.3%
weak226115-11150.9% 44.4% – 57.3%

Where those picks came from

MONEYLINE
 Good dataKnown-bad inputsEvery published pick
 recordhit %95% CIunitsrecordhit %95% CIunitsrecordhit %95% CIunits
all283-20058.6% 54.1% – 62.9%+13.6 61-4856.0% 46.6% – 64.9%+5.1 344-24858.1% 54.1% – 62.0%+18.7
HIGH24-1070.6% 53.8% – 83.2%+2.7 2-250.0% 15.0% – 85.0%-0.9 26-1268.4% 52.5% – 80.9%+1.7
MEDIUM108-5964.7% 57.2% – 71.5%+14.7 22-1756.4% 41.0% – 70.7%-1.9 130-7663.1% 56.3% – 69.4%+12.8
LOW151-13153.5% 47.7% – 59.3%-3.8 37-2956.1% 44.1% – 67.4%+7.9 188-16054.0% 48.8% – 59.2%+4.1
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other.

 nrecordhit %95% CIunitsb/e
RECONSTRUCTED · all21551138-101752.8% 50.7% – 54.9%-112.8 56.6%
backtest · all72053987-321855.3% 54.2% – 56.5%-334.1 58.2%
TOTAL* TESTED & FAILED — no edge over 8,549 past games.
 Good dataKnown-bad inputsEvery published pick
 recordhit %95% CIunitsrecordhit %95% CIunitsrecordhit %95% CIunits
all236-22850.9% 46.3% – 55.4%-15.3 29-3843.3% 32.1% – 55.2%-12.1 265-26649.9% 45.7% – 54.1%-27.4
strong47-3855.3% 44.7% – 65.4%+3.2 3-350.0% 18.8% – 81.2%-0.4 50-4154.9% 44.7% – 64.8%+2.8
mid47-4352.2% 42.0% – 62.2%-0.5 5-935.7% 16.3% – 61.2%-4.7 52-5250.0% 40.6% – 59.4%-5.2
weak142-14749.1% 43.4% – 54.9%-18.0 21-2644.7% 31.4% – 58.8%-7.0 163-17348.5% 43.2% – 53.8%-25.0
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other.

 nrecordhit %95% CIunitsb/e
RECONSTRUCTED · all1654845-80951.1% 48.7% – 53.5%-47.8 52.5%
backtest · all68953527-336851.2% 50.0% – 52.3%-285.8 53.4%
▸ Show why 40 totals picks (2026-08-10 to 2026-08-12) can never be settled

40 totals picks (2026-08-10 to 2026-08-12) can never be settled: they were published before the ledger recorded the posted line, and that line is not recoverable after the fact. They are not back-filled. Totals published from now on carry their line and settle normally.

NRFI* UNGRADED — no price history exists to check it against.
 Good dataKnown-bad inputsEvery published pick
 recordhit %95% CIrecordhit %95% CIrecordhit %95% CI
all224-25946.4% 42.0% – 50.8%70-3964.2% 54.9% – 72.6%294-29849.7% 45.6% – 53.7%
strong51-6544.0% 35.3% – 53.0%16-964.0% 44.5% – 79.8%67-7447.5% 39.5% – 55.7%
mid82-9745.8% 38.7% – 53.1%30-1665.2% 50.8% – 77.3%112-11349.8% 43.3% – 56.3%
weak91-9748.4% 41.4% – 55.5%24-1463.2% 47.3% – 76.6%115-11150.9% 44.4% – 57.3%
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other.

 nrecordhit %95% CI
RECONSTRUCTED · all21551150-100553.4% 51.3% – 55.5%
backtest · all72013712-348951.5% 50.4% – 52.7%

These three are never added together. Only LIVE is evidence.

▸ Show what LIVE, RECONSTRUCTED and backtest mean

LIVE — Picks this card actually published, saved before the games started. Read from data/live/published_picks.csv, which is written at publish time and refuses any pick for a game that has already started. It is the only source that can support a forward claim.

RECONSTRUCTED — A rebuilt history of 2026. These picks were never published — the model was re-run afterwards. NOT a live record. 2026 walk-forward output. Those picks came from a model fit on a different fold than the card uses, and every design choice in the pipeline was made with these results already visible. Treat as an upper bound, never as evidence.

backtest — 2021-2025, the years the model was designed on. 2021-2025 walk-forward. The model never saw these rows, but the feature pipeline was built while looking at this era.

RECONSTRUCTED postseason — hindsight, not a record
These picks were never published. They were produced after the games were played, and every choice behind them was made with the outcomes visible. They are not in the live, postseason, reconstructed-season or backtest records, and none of those contains any of these.
MarketW-L(-push)
Side7-2
Total*7-1-1
NRFI*5-4
27 row(s) over 3 slate(s): 2026-09-29, 2026-09-30, 2026-10-01. No hit rate, ROI, units or interval is shown, and none will be. A rate over a selection that could see the results measures the selection. 9 row(s) carry no price and therefore no break-even — no first-inning price was ever bought for these dates, and estimating one would be fabrication. Starters on every row were read from a schedule fetch made after the decision time, which is recorded per row as probables_source.
▸ How is it doing lately?

Every published pick, no data-quality filter — the same record the default view above shows. 1,715 settled over 46 slates.

Week over week
weekMONEYLINETOTAL*NRFI*
Aug 10–Aug 1648-44 52.2%22-29 43.1%60-32 65.2%
Aug 17–Aug 2360-34 63.8%40-48 45.5%45-49 47.9%
Aug 24–Aug 3028-19 59.6%22-24 47.8%24-23 51.1%
Aug 31–Sep 0638-47 44.7%44-39 53.0%39-46 45.9%
Sep 07–Sep 1355-36 60.4%41-43 48.8%37-54 40.7%
Sep 14–Sep 2066-28 70.2%47-44 51.6%45-49 47.9%
Sep 21–Sep 2749-40 55.1%49-39 55.7%44-45 49.4%

21 cell(s) here, so about 1.1 of them expected to look unusual by chance.

Month over month
monthMONEYLINETOTAL*NRFI*
August 2026136-97 58.4%84-101 45.4%129-104 55.4%
September 2026208-151 57.9%181-165 52.3%165-194 46.0%

6 cell(s) here, so about 0.3 of them expected to look unusual by chance.

By day of the week
dayMONEYLINETOTAL*NRFI*
Monday37-17 68.5%19-24 44.2%24-30 44.4%
Tuesday65-40 61.9%47-39 54.7%54-51 51.4%
Wednesday57-49 53.8%43-48 47.3%40-66 37.7%
Thursday30-23 56.6%21-28 42.9%32-21 60.4%
Friday61-37 62.2%48-46 51.1%45-53 45.9%
Saturday42-46 47.7%49-37 57.0%49-39 55.7%
Sunday52-36 59.1%38-44 46.3%50-38 56.8%

21 cell(s) here, so about 1.1 of them expected to look unusual by chance.

Are the stated probabilities still honest?

Mean stated probability minus what actually happened. Positive means the model was overconfident that period.

weekMONEYLINETOTAL*NRFI*
Aug 10+2.2
54.4% said / 52.2% got
+8.6
51.7% said / 43.1% got
-11.9
53.3% said / 65.2% got
Aug 17-8.8
55.1% said / 63.8% got
+7.0
52.5% said / 45.5% got
+5.6
53.5% said / 47.9% got
Aug 24-5.8
53.8% said / 59.6% got
+4.7
52.5% said / 47.8% got
+1.8
52.9% said / 51.1% got
Aug 31+9.3
54.0% said / 44.7% got
-0.8
52.2% said / 53.0% got
+7.6
53.5% said / 45.9% got
Sep 07-6.0
54.5% said / 60.4% got
+2.8
51.6% said / 48.8% got
+12.3
53.0% said / 40.7% got
Sep 14-16.0
54.2% said / 70.2% got
+1.5
53.1% said / 51.6% got
+4.9
52.8% said / 47.9% got
Sep 21-0.4
54.7% said / 55.1% got
-0.5
55.2% said / 55.7% got
+3.9
53.4% said / 49.4% got

21 cell(s), so about 1.1 expected to look unusual by chance.

Rolling 100-pick window

Hit rate over the last 100 picks at every point, in date order. Calendar weeks are an arbitrary cut; this has no boundaries to straddle.

MONEYLINE — now 56.0%, range 45.0% to 71.0% across 493 windows
TOTAL* — now 54.0%, range 42.0% to 56.0% across 432 windows
NRFI* — now 48.0%, range 39.0% to 66.0% across 493 windows

3 series shown.

Cumulative units, against the base rate at the same prices

Solid line: your picks. Dashed: what a bettor hitting exactly the base rate would have earned at the prices you actually paid. NRFI is absent because no NRFI price exists — units for it would have to be invented.

MONEYLINE — you +18.7u, base rate -22.1u, difference +40.7u over 571 priced picks
TOTAL* — you -27.4u, base rate -16.0u, difference -11.5u over 531 priced picks

2 series shown; about 0.1 expected to look unusual by chance.

The honest headline

Tested on 1,729 real 2026 games at real prices, this lost 6.5%.

That is what paying the bookmaker’s cut (the “vig”) costs you, with no skill at all.

Nothing here has been shown to make money.

▸ What to trust, and what the words mean

What to trust

Side is the only column checked against real results — and it came out in the right order.

Total and NRFI have not been checked. The model has an opinion; nobody has verified it.

What the confidence words mean

HIGH / MED / LOW on Side were measured against what actually happened.

weak / mid / strong on Total and NRFI are only how sure the model feels. Untested.

Different words on purpose, so one can never be read as the other.

▸ Full technical detail

TOTAL* — tested and FAILED. Measured on 8,549 out-of-sample games: HIGH 53.5% [51.0%, 56.0%] n=1,528, MEDIUM 52.1% [50.3%, 54.0%] n=2,868, LOW 49.6% [48.1%, 51.1%] n=4,153. The ordering holds, but the intervals for HIGH/MEDIUM and MEDIUM/LOW overlap, so those are not shown to be distinct. HIGH's interval clears 50% by only 0.96 points. Incremental R-squared over the posted line is +0.0022 (measured 2026-08-09, 8,272 games). Calibrated but uninformative.

NRFI* — UNGRADED, never tested at all. No historical NRFI price exists anywhere, so hit-rate-against-break-even and Brier-against-market are uncomputable, not merely poor. No break-even is shown because there is no price.

weak / mid / strong is model conviction, NOT validation. It is which third of the claim-strength distribution the number falls in, and has not been shown to predict anything. A “strong” total is not a HIGH side.

SIDE — the only validated column. Measured on 9,360 out-of-sample games: HIGH 58.5% [56.2%, 60.9%] n=1,708, MEDIUM 56.9% [55.1%, 58.7%] n=2,911, LOW 52.1% [50.6%, 53.5%] n=4,741. The ordering holds, but the intervals for HIGH/MEDIUM overlap, so those are not shown to be distinct. MEDIUM over LOW is clean. Only side HIGH is shaded. Break-even is shown beside every priced pick and never filters — every game gets an output, there is no “no bet”.

When Total* is empty there are two reasons. Either no posted total exists, in which case the model has nothing to read P(over) off and the blank is a missing input; or a line exists but the game is inside the 21-day sealed window that is deliberately not graded. Which one applies is stated on the game itself, and nowhere else — this note used to end by naming the sealed window as the usual cause, which contradicted the banner at the top of the page whenever the real cause was a missing line. Two explanations for one blank is worse than none: it sends the reader to fix the wrong thing.

Recalibration. Probabilities are recalibrated against the trailing year, because the raw model runs overconfident on recent data: reliability 0.0008 on the original window against 0.0026 on 2026, always in the same direction. The correction is a Platt map fitted strictly earlier than the slate it is applied to, and it is monotone, so band ordering is untouched — only the numbers move. On 2026 it cuts reliability to 0.0021 and narrows the HIGH band's stated-minus-actual gap from +0.0450 to +0.0076 — and closes it, on this window. On the late-2025 tail it makes it WORSE (0.0032 to 0.0039, n=838). Measured over the 6,914 rows that can be calibrated at all.

Notices and system health

Standing explanations. Nothing here needs a decision.

▸ Show what POSTSEASON means for these picks (AL Division Series, NL Division Series)
POSTSEASON SLATE — AL Division Series, NL Division Series
The model has never seen a postseason game; the methodology is identical to the regular season.
▸ What this means for these picks
The model has never seen a postseason game. Its history is regular-season only, so these picks come from a model fitted on a different population from the one it is predicting — rotations run on short rest, bullpens are used at leverage rather than by role, rosters are reshaped, and the field is a selected group of clubs rather than the whole league.

The methodology is identical to the regular season. Same features, same fit window, same regularisation, same bands. Nothing in the model varies by game type — which is the point: it cannot tell, so this notice has to.

The band on a postseason pick is the model’s conviction, not a measured hit rate. Every band figure, interval and hit rate quoted anywhere on this page was measured on regular-season games and says nothing about these.

The base rates the tables colour against are regular-season measurements. Backing the home team every day wins at a rate measured over regular-season games; postseason home-field advantage is a different number and nobody here has measured it.

These picks are recorded in a separate ledger (postseason_picks.csv) and are never pooled with the live, reconstructed or backtest records. Until a postseason sample exists the card shows that record and no rate.
▸ DECISION TIME WAS 15:30 UTC (08:30 Phoenix)

DECISION TIME WAS 15:30 UTC (08:30 Phoenix) — but this card was NOT priced then. The snapshot behind it was collected at 20:00Z, 4h 30m late. Prices, break-evens and edges below are as of 20:00Z.

▸ Show API credits: 19,808 left

API 19,808 credits left of 20,000 (99% left) · as of this build · nrfi · resets in 25 days · ~15/day · ~19,400 left at reset · 17,400 spare for a one-off purchase

▸ Show why 330 published picks were built from inputs known to be wrong

330 of the published picks were built from inputs known to be wrong. The buttons in the season block above include or exclude them, and the two groups are never pooled into one rate.

▸ Show why -6.5% ROI over 1,729 priced 2026 games is what paying the vig looks like

2026 backtest at real prices: -6.5% ROI over 1,729 priced games. The loss was statistically indistinguishable from paying the vig with zero skill (residual -2.6%, z=-1.19).

How the model has actually done

Neither section below is a live record — both are the model re-run over past games. The live record is at the top of this page. Three markets, never pooled, never added together.

▸ Show 2021–2025 — the years the model was designed on

2021–2025. The model never saw these games, but the features were designed while looking at this era.

Moneyline validated — but HIGH and MEDIUM no longer separate
Winning? No — -334.1 units over 7,192 priced picks, a -4.6% return at the prices actually paid. Win rate 55.3%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly honest — it claimed 56.8% and delivered 55.3%, 1.5 points too high. Worst on its HIGH picks: claimed 62.5%, got 58.9%.
Bands separating? They ORDER but do not SEPARATE — HIGH 58.9% down to LOW 52.6%, but the intervals for HIGH and MEDIUM overlap, so those are not distinguishable.
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
HIGH1,38058.9%56.3% – 61.5%62.5%+3.5 pts-88.8-6.4%
MEDIUM2,32657.4%55.4% – 59.4%58.2%+0.8 pts-73.9-3.2%
LOW3,49952.6%50.9% – 54.2%53.7%+1.2 pts-171.4-4.9%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
16.5%-44.0%90139.4%43.0%-3.5 pts
44.1%-47.9%90046.2%47.3%-1.2 pts
47.9%-50.8%90149.4%51.1%-1.6 pts
50.8%-53.4%90052.1%50.1%+2.0 pts
53.5%-55.7%90154.6%57.7%-3.1 pts
55.7%-58.3%90056.9%54.6%+2.4 pts
58.3%-62.0%90160.0%59.4%+0.6 pts
62.0%-79.8%90165.9%63.3%+2.6 pts
Total* TESTED AND FAILED — bands measured on 8,272 games and did not separate
Winning? No — -285.8 units over 6,895 priced picks, a -4.1% return at the prices actually paid. Win rate 51.2%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly overconfident — it claimed 56.3% and delivered 51.2%, 5.1 points too high. Worst on its strong picks: claimed 61.9%, got 54.0%.
Bands separating? They ORDER but do not SEPARATE — strong 54.0% down to weak 47.9%, but the intervals for strong and mid, mid and weak overlap, so those are not distinguishable.
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
strong2,29954.0%51.9% – 56.0%61.9%+7.9 pts+20.80.9%
mid2,29851.6%49.5% – 53.6%55.3%+3.7 pts-77.5-3.4%
weak2,29847.9%45.9% – 50.0%51.7%+3.8 pts-229.2-10.0%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
6.4%-38.6%86234.6%46.6%-12.0 pts
38.6%-42.2%86240.6%46.2%-5.6 pts
42.2%-44.8%86243.5%45.1%-1.6 pts
44.8%-47.1%86145.9%51.0%-5.1 pts
47.1%-49.3%86248.2%52.2%-4.0 pts
49.3%-52.0%86250.6%48.4%+2.2 pts
52.0%-55.3%86253.5%46.9%+6.6 pts
55.3%-79.8%86258.9%54.4%+4.4 pts
NRFI* UNGRADED — no price history was ever collected, so there is nothing to grade against
Winning? Cannot be answered. No historical price exists for this market, so there is no break-even to measure against — win rate is 51.5% over 7,201 picks, and a perfectly calibrated 52% laid at -125 loses all season.
Calibrated? Broadly overconfident — it claimed 53.8% and delivered 51.5%, 2.2 points too high. Worst on its strong picks: claimed 57.1%, got 52.4%.
Bands separating? No — they do not even order: weak is beating mid. A strength label that runs backwards is carrying no information.
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff by
strong2,40152.4%50.4% – 54.4%57.1%+4.7 pts
mid2,40050.9%48.9% – 52.9%53.2%+2.3 pts
weak2,40051.3%49.3% – 53.3%51.0%-0.3 pts
▸ Show 2026 rebuilt after the fact — an upper bound, never a record

2026, rebuilt afterwards. These picks were never published — an upper bound, never evidence.

Moneyline validated — but HIGH and MEDIUM no longer separate
Winning? No — -112.8 units over 1,729 priced picks, a -6.5% return at the prices actually paid. Win rate 52.8%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly overconfident — it claimed 56.3% and delivered 52.8%, 3.5 points too high. Worst on its HIGH picks: claimed 61.5%, got 57.0%.
Bands separating? They ORDER but do not SEPARATE — HIGH 57.0% down to LOW 50.6%, but the intervals for HIGH and MEDIUM, MEDIUM and LOW overlap, so those are not distinguishable.
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
HIGH32857.0%51.6% – 62.3%61.5%+4.5 pts-18.8-5.7%
MEDIUM58555.0%51.0% – 59.0%58.5%+3.4 pts-26.6-4.7%
LOW1,24250.6%47.9% – 53.4%54.0%+3.3 pts-67.4-8.1%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
14.8%-46.5%27042.0%45.2%-3.2 pts
46.5%-50.1%26948.6%55.0%-6.5 pts
50.1%-52.4%26951.3%48.0%+3.4 pts
52.4%-54.3%26953.4%49.8%+3.6 pts
54.3%-56.0%27055.1%50.4%+4.8 pts
56.1%-58.2%26957.1%54.3%+2.8 pts
58.2%-61.4%26959.7%58.4%+1.4 pts
61.4%-79.5%27064.6%59.6%+5.0 pts
Total* TESTED AND FAILED — bands measured on 8,272 games and did not separate
Winning? No — -47.8 units over 1,652 priced picks, a -2.9% return at the prices actually paid. Win rate 51.1%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly overconfident — it claimed 55.9% and delivered 51.1%, 4.8 points too high. Worst on its strong picks: claimed 61.2%, got 51.4%.
Bands separating? They ORDER but do not SEPARATE — strong 51.4% down to weak 50.8%, but the intervals for strong and mid, mid and weak overlap, so those are not distinguishable.
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
strong55251.4%47.3% – 55.6%61.2%+9.8 pts-13.5-2.5%
mid55151.0%46.8% – 55.2%54.9%+3.9 pts-18.3-3.3%
weak55150.8%46.7% – 55.0%51.5%+0.7 pts-16.0-2.9%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
8.6%-39.6%20734.9%50.7%-15.8 pts
39.6%-42.8%20741.2%52.2%-10.9 pts
42.8%-45.1%20643.9%47.6%-3.6 pts
45.1%-47.2%20746.0%46.9%-0.8 pts
47.2%-49.3%20748.2%49.8%-1.5 pts
49.3%-51.6%20650.4%49.5%+0.9 pts
51.7%-54.7%20753.0%44.9%+8.0 pts
54.7%-67.2%20757.8%55.1%+2.7 pts
NRFI* UNGRADED — no price history was ever collected, so there is nothing to grade against
Winning? Cannot be answered. No historical price exists for this market, so there is no break-even to measure against — win rate is 53.4% over 2,155 picks, and a perfectly calibrated 52% laid at -125 loses all season.
Calibrated? Broadly honest — it claimed 54.6% and delivered 53.4%, 1.2 points too high. Worst on its strong picks: claimed 58.8%, got 53.5%.
Bands separating? They ORDER but do not SEPARATE — strong 53.5% down to weak 53.2%, but the intervals for strong and mid, mid and weak overlap, so those are not distinguishable.
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff by
strong71953.5%49.9% – 57.2%58.8%+5.3 pts
mid71853.3%49.7% – 57.0%53.7%+0.4 pts
weak71853.2%49.5% – 56.8%51.1%-2.1 pts
▸ Show the full technical report (fixed-width, as the terminal prints it)
====================================================================================================
TRACKING — THREE SEPARATE LEDGERS
====================================================================================================
These three markets are NEVER pooled. Different base rates, different
validation status, different meanings. There is no combined record
anywhere in this report, by design.

Live begins 2026-08-10 — the first slate a card actually published, read from the ledger. Backtest and live are never summed.

####################################################################################################
# BACKTEST
# 2021-2025 walk-forward. The model never saw these rows, but the
# feature pipeline was designed while looking at this era.
####################################################################################################

----------------------------------------------------------------------------------------------------
MONEYLINE — the only validated market. The band claim below is
            measured on this run, not quoted from memory.
            Measured on 9,360 games: HIGH 58.5% / MEDIUM 56.9% / LOW 52.1%.
            They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM overlap.
            Clean: MEDIUM over LOW.
            Intervals: HIGH [56.2%,60.9%] n=1,708, MEDIUM [55.1%,58.7%] n=2,911, LOW [50.6%,53.5%] n=4,741.
----------------------------------------------------------------------------------------------------
  hit rate by band, with 95% CI, stated vs actual, and units.
  n = all graded games (hit rate, calibration).
  n_priced = the subset with a real price (units, ROI). The 2026
  archive backfill priced 2,299 of the 2,301 games that had none,
  so n_priced now nearly equals n; the two games still missing
  start before the snapshot that was bought.
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual   units   roi  n_priced
      HIGH 1380    58.9% 56.3% 61.5%     5.2%  62.5%            +3.5 pts  -88.81 -6.4%      1380
    MEDIUM 2326    57.4% 55.4% 59.4%     4.0%  58.2%            +0.8 pts  -73.90 -3.2%      2321
       LOW 3499    52.6% 50.9% 54.2%     3.3%  53.7%            +1.2 pts -171.38 -4.9%      3491

  calibration by probability bucket:
         bucket   n stated actual      gap ci_lo ci_hi
    16.5%-44.0% 901  39.4%  43.0% -3.5 pts 39.8% 46.2%
    44.1%-47.9% 900  46.2%  47.3% -1.2 pts 44.1% 50.6%
    47.9%-50.8% 901  49.4%  51.1% -1.6 pts 47.8% 54.3%
    50.8%-53.4% 900  52.1%  50.1% +2.0 pts 46.9% 53.4%
    53.5%-55.7% 901  54.6%  57.7% -3.1 pts 54.5% 60.9%
    55.7%-58.3% 900  56.9%  54.6% +2.4 pts 51.3% 57.8%
    58.3%-62.0% 901  60.0%  59.4% +0.6 pts 56.1% 62.5%
    62.0%-79.8% 901  65.9%  63.3% +2.6 pts 60.1% 66.3%

----------------------------------------------------------------------------------------------------
TOTALS* — BANDS WERE TESTED AND FAILED. Incremental R^2 over the
          posted line is +0.0022, and weak/mid/strong is CLAIM
          STRENGTH ONLY, never a validated band.
          Measured on 8,549 games: HIGH 53.5% / MEDIUM 52.1% / LOW 49.6%.
          They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM, MEDIUM and LOW overlap.
          Intervals: HIGH [51.0%,56.0%] n=1,528, MEDIUM [50.3%,54.0%] n=2,868, LOW [48.1%,51.1%] n=4,153.
----------------------------------------------------------------------------------------------------
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual   units    roi  n_priced
    strong 2299    54.0% 51.9% 56.0%     4.1%  61.9%            +7.9 pts   20.81   0.9%      2299
       mid 2298    51.6% 49.5% 53.6%     4.1%  55.3%            +3.7 pts  -77.46  -3.4%      2298
      weak 2298    47.9% 45.9% 50.0%     4.1%  51.7%            +3.8 pts -229.17 -10.0%      2298

  calibration by probability bucket:
         bucket   n stated actual       gap ci_lo ci_hi
     6.4%-38.6% 862  34.6%  46.6% -12.0 pts 43.3% 50.0%
    38.6%-42.2% 862  40.6%  46.2%  -5.6 pts 42.9% 49.5%
    42.2%-44.8% 862  43.5%  45.1%  -1.6 pts 41.8% 48.5%
    44.8%-47.1% 861  45.9%  51.0%  -5.1 pts 47.7% 54.3%
    47.1%-49.3% 862  48.2%  52.2%  -4.0 pts 48.9% 55.5%
    49.3%-52.0% 862  50.6%  48.4%  +2.2 pts 45.1% 51.7%
    52.0%-55.3% 862  53.5%  46.9%  +6.6 pts 43.6% 50.2%
    55.3%-79.8% 862  58.9%  54.4%  +4.4 pts 51.1% 57.7%

----------------------------------------------------------------------------------------------------
NRFI* — UNGRADED. No historical price is HELD for this market, so
        break-even, units and market comparison are UNCOMPUTABLE,
        not merely absent. Hit rate only.
        (A season of history is purchasable at ~17,210 credits and
         has not been bought. The LIVE market is obtainable:
         DraftKings quoted the 0.5 line on 3 of the 3 events
         carrying it on 2026-08-13, under the market key
         alternate_totals_1st_1_innings.)
----------------------------------------------------------------------------------------------------
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual
    strong 2401    52.4% 50.4% 54.4%     4.0%  57.1%            +4.7 pts
       mid 2400    50.9% 48.9% 52.9%     4.0%  53.2%            +2.3 pts
      weak 2400    51.3% 49.3% 53.3%     4.0%  51.0%            -0.3 pts

####################################################################################################
# RECONSTRUCTED
# 2026 REBUILT FROM WALK-FORWARD OUTPUT. NOT a live record: these
# picks were never published, the model was re-run afterwards on a
# different fold, and the pipeline was designed with these results
# already visible. An upper bound, never evidence.
# The live record lives in data/live/published_picks.csv and is shown on the card.
####################################################################################################

----------------------------------------------------------------------------------------------------
MONEYLINE — the only validated market. The band claim below is
            measured on this run, not quoted from memory.
            Measured on 9,360 games: HIGH 58.5% / MEDIUM 56.9% / LOW 52.1%.
            They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM overlap.
            Clean: MEDIUM over LOW.
            Intervals: HIGH [56.2%,60.9%] n=1,708, MEDIUM [55.1%,58.7%] n=2,911, LOW [50.6%,53.5%] n=4,741.
----------------------------------------------------------------------------------------------------
  hit rate by band, with 95% CI, stated vs actual, and units.
  n = all graded games (hit rate, calibration).
  n_priced = the subset with a real price (units, ROI). The 2026
  archive backfill priced 2,299 of the 2,301 games that had none,
  so n_priced now nearly equals n; the two games still missing
  start before the snapshot that was bought.
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual  units   roi  n_priced
      HIGH  328    57.0% 51.6% 62.3%    10.7%  61.5%            +4.5 pts -18.79 -5.7%       327
    MEDIUM  585    55.0% 51.0% 59.0%     8.0%  58.5%            +3.4 pts -26.62 -4.7%       572
       LOW 1242    50.6% 47.9% 53.4%     5.6%  54.0%            +3.3 pts -67.43 -8.1%       830

  calibration by probability bucket:
         bucket   n stated actual      gap ci_lo ci_hi
    14.8%-46.5% 270  42.0%  45.2% -3.2 pts 39.4% 51.1%
    46.5%-50.1% 269  48.6%  55.0% -6.5 pts 49.0% 60.9%
    50.1%-52.4% 269  51.3%  48.0% +3.4 pts 42.1% 53.9%
    52.4%-54.3% 269  53.4%  49.8% +3.6 pts 43.9% 55.7%
    54.3%-56.0% 270  55.1%  50.4% +4.8 pts 44.4% 56.3%
    56.1%-58.2% 269  57.1%  54.3% +2.8 pts 48.3% 60.1%
    58.2%-61.4% 269  59.7%  58.4% +1.4 pts 52.4% 64.1%
    61.4%-79.5% 270  64.6%  59.6% +5.0 pts 53.7% 65.3%

----------------------------------------------------------------------------------------------------
TOTALS* — BANDS WERE TESTED AND FAILED. Incremental R^2 over the
          posted line is +0.0022, and weak/mid/strong is CLAIM
          STRENGTH ONLY, never a validated band.
          Measured on 8,549 games: HIGH 53.5% / MEDIUM 52.1% / LOW 49.6%.
          They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM, MEDIUM and LOW overlap.
          Intervals: HIGH [51.0%,56.0%] n=1,528, MEDIUM [50.3%,54.0%] n=2,868, LOW [48.1%,51.1%] n=4,153.
----------------------------------------------------------------------------------------------------
     group   n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual  units   roi  n_priced
    strong 552    51.4% 47.3% 55.6%     8.3%  61.2%            +9.8 pts -13.54 -2.5%       552
       mid 551    51.0% 46.8% 55.2%     8.3%  54.9%            +3.9 pts -18.25 -3.3%       549
      weak 551    50.8% 46.7% 55.0%     8.3%  51.5%            +0.7 pts -16.03 -2.9%       551

  calibration by probability bucket:
         bucket   n stated actual       gap ci_lo ci_hi
     8.6%-39.6% 207  34.9%  50.7% -15.8 pts 44.0% 57.5%
    39.6%-42.8% 207  41.2%  52.2% -10.9 pts 45.4% 58.9%
    42.8%-45.1% 206  43.9%  47.6%  -3.6 pts 40.9% 54.4%
    45.1%-47.2% 207  46.0%  46.9%  -0.8 pts 40.2% 53.7%
    47.2%-49.3% 207  48.2%  49.8%  -1.5 pts 43.0% 56.5%
    49.3%-51.6% 206  50.4%  49.5%  +0.9 pts 42.8% 56.3%
    51.7%-54.7% 207  53.0%  44.9%  +8.0 pts 38.3% 51.7%
    54.7%-67.2% 207  57.8%  55.1%  +2.7 pts 48.3% 61.7%

----------------------------------------------------------------------------------------------------
NRFI* — UNGRADED. No historical price is HELD for this market, so
        break-even, units and market comparison are UNCOMPUTABLE,
        not merely absent. Hit rate only.
        (A season of history is purchasable at ~17,210 credits and
         has not been bought. The LIVE market is obtainable:
         DraftKings quoted the 0.5 line on 3 of the 3 events
         carrying it on 2026-08-13, under the market key
         alternate_totals_1st_1_innings.)
----------------------------------------------------------------------------------------------------
     group   n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual
    strong 719    53.5% 49.9% 57.2%     7.3%  58.8%            +5.3 pts
       mid 718    53.3% 49.7% 57.0%     7.3%  53.7%            +0.4 pts
      weak 718    53.2% 49.5% 56.8%     7.3%  51.1%            -2.1 pts

====================================================================================================
READING THESE NUMBERS
====================================================================================================
  Every hit rate above carries a 95% interval. A rate whose interval
  contains the relevant break-even (or 0.50) has not demonstrated
  anything, however good the point estimate looks.
  Live moneyline has 483 settled picks over 36 slates (about 13 a slate), and at that count the 95% interval is +/- 4.4 points.
  Narrowing it to +/- 3 points — the width at which a live rate could be told apart from the no-model base rate — needs roughly 1,036 settled picks, about 41 more slates at the current rate.
  Live conclusions should not be drawn before then.