MLB card — 2026-08-21

15 games · times in Phoenix · generated 2026-08-21 14:44 UTC
Side is the only validated column. Total* and NRFI* are not — see the notes at the bottom.
Probabilities: calibrated on 1,968 games from the prior year.

DECISION TIME 15:30 UTC (08:30 Phoenix) — every price, probability, break-even and edge on this card is as of that moment, not as of now.

API 15,987 of 20,000 credits (80%) · as of this build · nrfi · resets in 11 days · ~48/day · ~15,400 left at reset · 13,400 spare for a one-off purchase

Season to datethree markets, never pooled

2026 backtest at real prices: -7.0% ROI over 1,443 priced games. The loss was statistically indistinguishable from paying the vig with zero skill (residual -3.0%, z=-1.26).

Showing clean picks only. 309 earlier picks were built from stale data and are counted separately — switch below.

Good data (Aug 18 →)Stale data (Aug 10–17)Compare both
LIVE since 2026-08-18
MONEYLINE
▸ Show the measurements

Measured on 8,650 out-of-sample games: HIGH 59.4% [57.1%, 61.8%] n=1,657, MEDIUM 57.1% [55.2%, 58.9%] n=2,772, LOW 52.2% [50.7%, 53.7%] n=4,221. The ordering holds, but the intervals for HIGH/MEDIUM overlap, so those are not shown to be distinct. MEDIUM over LOW is clean.

 nrecordhit %95% CIunitsb/e
all3920-1951.3%36.2% – 66.1%-5.056.3%
HIGH21-1too few-0.463.1%
MEDIUM148-6too few-1.060.0%
LOW2311-12too few-3.653.5%
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other. RECONSTRUCTED is 2026 rebuilt after the fact — an upper bound, not a record. Backtest is the era the features were designed against.

 nrecordhit %95% CIunitsb/e
RECONSTRUCTED · all1445758-68752.5%49.9% – 55.0%-99.356.5%
backtest · all72054014-319155.7%54.6% – 56.9%-293.858.3%
TOTAL* Tested on 8,272 past games and it did not work — no proven edge.
▸ Show the measurements

Bands were tested out-of-sample and did NOT separate, so no band is shown on this column — showing one would be decoration. Measured on 8,272 out-of-sample games: HIGH 52.5% [50.0%, 55.1%] n=1,513, MEDIUM 52.5% [50.7%, 54.4%] n=2,754, LOW 50.1% [48.6%, 51.7%] n=4,005. The ordering holds, but the intervals for HIGH/MEDIUM and MEDIUM/LOW overlap, so those are not shown to be distinct. HIGH's interval clears 50% by only 0.03 points. Incremental R-squared over the posted line is +0.0022 (measured 2026-08-09, 8,272 games). Calibrated but uninformative.

 nrecordhit %95% CIunitsb/e
all3718-1948.6%33.4% – 64.1%-2.452.3%
strong51-4too few-3.253.0%
mid52-3too few-1.251.9%
weak2715-12too few+1.952.3%
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other. RECONSTRUCTED is 2026 rebuilt after the fact — an upper bound, not a record. Backtest is the era the features were designed against.

 nrecordhit %95% CIunitsb/e
RECONSTRUCTED · all1377703-67451.1%48.4% – 53.7%-41.452.5%
backtest · all68953546-334951.4%50.2% – 52.6%-250.053.4%
NRFI* Never checked — no price history was kept to check it against.
▸ Show the measurements

No historical NRFI price was ever collected, so hit-rate-against-break-even and Brier-against-market are UNCOMPUTABLE for any past game — uncomputable, not merely poor. That is the whole reason this column is ungraded, and it is a purchase decision: a season of first-inning history costs about 17,210 credits against a 20,000/month plan, and has not been judged worth it. The market itself is live and obtainable. Measured 2026-08-13 across the three events on the feed carrying any first-inning market, DraftKings quoted the 0.5 line on 3 of 3 — Over 0.5 -125 / Under 0.5 +100 and similar — under the market key 'alternate_totals_1st_1_innings'. Fanatics carried the same key on 3 of 3 but only at 1.5, which is a different bet. Earlier wording here said no price existed at either book; that was an artefact of asking for the wrong market key, not a fact about the board.

 nrecordhit %95% CI
all3916-2341.0%27.1% – 56.6%
strong72-5too few
mid199-10too few
weak135-8too few
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other. RECONSTRUCTED is 2026 rebuilt after the fact — an upper bound, not a record. Backtest is the era the features were designed against.

 nrecordhit %95% CI
RECONSTRUCTED · all1445768-67753.1%50.6% – 55.7%
backtest · all72013712-348951.5%50.4% – 52.7%

These three are never added together. Only LIVE is evidence.

▸ Show what LIVE, RECONSTRUCTED and backtest mean

LIVE — Picks this card actually published, saved before the games started. Read from data/live/published_picks.csv, which is written at publish time and refuses any pick for a game that has already started. It is the only source that can support a forward claim.

RECONSTRUCTED — A rebuilt history of 2026. These picks were never published — the model was re-run afterwards. NOT a live record. 2026 walk-forward output. Those picks came from a model fit on a different fold than the card uses, and every design choice in the pipeline was made with these results already visible. Treat as an upper bound, never as evidence.

backtest — 2021-2025, the years the model was designed on. 2021-2025 walk-forward. The model never saw these rows, but the feature pipeline was built while looking at this era.

The honest headline

Tested on 1,443 real 2026 games at real prices, this lost 6.1%.

That is what paying the bookmaker’s cut (the “vig”) costs you, with no skill at all.

Nothing here has been shown to make money.

▸ What to trust, and what the words mean

What to trust

Side is the only column checked against real results — and it came out in the right order.

Total and NRFI have not been checked. The model has an opinion; nobody has verified it.

What the confidence words mean

HIGH / MED / LOW on Side were measured against what actually happened.

weak / mid / strong on Total and NRFI are only how sure the model feels. Untested.

Different words on purpose, so one can never be read as the other.

▸ Full technical detail

TOTAL* — tested and FAILED. Measured on 8,272 out-of-sample games: HIGH 52.5% [50.0%, 55.1%] n=1,513, MEDIUM 52.5% [50.7%, 54.4%] n=2,754, LOW 50.1% [48.6%, 51.7%] n=4,005. The ordering holds, but the intervals for HIGH/MEDIUM and MEDIUM/LOW overlap, so those are not shown to be distinct. HIGH's interval clears 50% by only 0.03 points. Incremental R-squared over the posted line is +0.0022 (measured 2026-08-09, 8,272 games). Calibrated but uninformative.

NRFI* — UNGRADED, never tested at all. No historical NRFI price exists anywhere, so hit-rate-against-break-even and Brier-against-market are uncomputable, not merely poor. No break-even is shown because there is no price.

weak / mid / strong is model conviction, NOT validation. It is which third of the claim-strength distribution the number falls in, and has not been shown to predict anything. A “strong” total is not a HIGH side.

SIDE — the only validated column. Measured on 8,650 out-of-sample games: HIGH 59.4% [57.1%, 61.8%] n=1,657, MEDIUM 57.1% [55.2%, 58.9%] n=2,772, LOW 52.2% [50.7%, 53.7%] n=4,221. The ordering holds, but the intervals for HIGH/MEDIUM overlap, so those are not shown to be distinct. MEDIUM over LOW is clean. Only side HIGH is shaded. Break-even is shown beside every priced pick and never filters — every game gets an output, there is no “no bet”.

When Total* is empty there are two reasons. Either no posted total exists, in which case the model has nothing to read P(over) off and the blank is a missing input; or a line exists but the game is inside the 21-day sealed window that is deliberately not graded. Since the archive backfill posted lines for 2,299 games of the 2026 season, the second is the usual reason on the newest slate.

Probabilities are recalibrated. Probabilities are recalibrated against the trailing year, because the raw model runs overconfident on recent data: reliability 0.0008 on the original window against 0.0039 on 2026, always in the same direction. The correction is a Platt map fitted strictly earlier than the slate it is applied to, and it is monotone, so band ordering is untouched — only the numbers move. On 2026 it cuts reliability to 0.0022 and narrows the HIGH band's stated-minus-actual gap from +0.0867 to +0.0517 — a real improvement that does NOT close it. On the late-2025 tail it improves it (0.0037 to 0.0026, n=838). Measured over the 6,204 rows that can be calibrated at all.

Has this been proven yet? ▸ show the workings
SIDE: behind, but too few picks to tell — 39 of 200 picks
TOTAL: behind, but too few picks to tell — 37 of 200 picks
NRFI: behind, but too few picks to tell — 39 of 200 picks

A win rate only means something when there are enough picks AND it beats what you would get with no model at all — always backing the home team, always the under, always NRFI. Both conditions, per market, per band. Nothing is hidden; what changes is whether a number is called reliable.

“Cannot clear” has to clear the same sample bar as “proven”. A tier that is behind on a handful of picks is having a bad run, not failing, so a negative verdict needs as many picks as a positive one — below that bar the row reads behind, too few picks to tell and nothing stronger.

▸ Why there are two groups of picks

Some picks were built while a data file was out of date, so the model was working from stale information. They are still shown and still counted in the record — they are just kept separate, because mixing them in would hide whether we proved the model or proved the bug. 2026-08-10 to 2026-08-17: pitcher_gamelogs_full.parquet was stale (last updated 2026-08-06). 24 of the model's 54 features derive from that file, so published probabilities were computed from starter form 4-11 days out of date. Measured on the 2026-08-17 slate: 1 of 11 moneyline picks flipped side, 5 of 11 bands changed, mean move 0.91 points. Totals moved most (mean 3.46 pts); NRFI least (0 of 11 flipped, mean 0.34 pts).

 recordsettledneededverdict
SIDE20-1939200behind, too few picks to tell
  HIGH1-12200behind, too few picks to tell
  MEDIUM8-614200too few picks to tell
  LOW11-1223200behind, too few picks to tell
TOTAL18-1937200behind, too few picks to tell
  strong1-45200behind, too few picks to tell
  mid2-35200behind, too few picks to tell
  weak15-1227200too few picks to tell
NRFI16-2339200behind, too few picks to tell
  strong2-57200behind, too few picks to tell
  mid9-1019200behind, too few picks to tell
  weak5-813200behind, too few picks to tell

NRFI disagrees with itself. The live record says 41.0% (16-23) on 39 settled picks — far worse than the same model rebuilt over the 2026 season, which says 51.8% on 8,646 picks, a sample 221 times larger. Both are shown because one of them is wrong and the small one is far more likely to be the misleading one: the larger sample is the one built from more evidence. Do not read the live figure as a finding.

How the model has actually done

Neither section below is a live record — both are the model re-run over past games. The live record is at the top of this page. Three markets, never pooled, never added together.

▸ Show 2021–2025 — the years the model was designed on

2021–2025. The model never saw these games, but the features were designed while looking at this era.

Moneyline validated — but HIGH and MEDIUM no longer separate
Winning? No — -293.8 units over 7,192 priced picks, a -4.1% return at the prices actually paid. Win rate 55.7%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly honest — it claimed 56.7% and delivered 55.7%, 1.0 points too high. Worst on its HIGH picks: claimed 62.4%, got 60.7%.
Bands separating? They ORDER but do not SEPARATE — HIGH 60.7% down to LOW 53.2%, but the intervals for HIGH and MEDIUM, MEDIUM and LOW overlap, so those are not distinguishable.
Big enough to say? Yes — 7,205 picks, interval 2.3 points wide (54.6% to 56.9%).
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
HIGH1,39060.7%58.1% – 63.3%62.4%+1.7 pts-54.2-3.9%
MEDIUM2,27856.6%54.5% – 58.6%58.0%+1.4 pts-96.2-4.2%
LOW3,53753.2%51.5% – 54.8%53.7%+0.5 pts-143.4-4.1%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
13.6%-44.2%90139.5%44.2%-4.6 pts
44.2%-48.1%90046.3%46.1%+0.2 pts
48.1%-50.8%90149.5%52.6%-3.1 pts
50.8%-53.3%90052.1%51.0%+1.1 pts
53.3%-55.6%90154.4%54.4%+0.0 pts
55.6%-58.2%90056.8%56.7%+0.2 pts
58.2%-61.8%90159.8%58.9%+0.9 pts
61.8%-80.9%90165.8%62.5%+3.3 pts
Total* TESTED AND FAILED — bands measured on 8,272 games and did not separate
Winning? No — -250.0 units over 6,895 priced picks, a -3.6% return at the prices actually paid. Win rate 51.4%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly overconfident — it claimed 56.3% and delivered 51.4%, 4.9 points too high. Worst on its strong picks: claimed 61.9%, got 54.1%.
Bands separating? They ORDER but do not SEPARATE — strong 54.1% down to weak 49.2%, but the intervals for strong and mid, mid and weak overlap, so those are not distinguishable.
Big enough to say? Yes — 6,895 picks, interval 2.4 points wide (50.2% to 52.6%).
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
strong2,29954.1%52.1% – 56.1%61.9%+7.8 pts+25.91.1%
mid2,29851.0%49.0% – 53.0%55.3%+4.3 pts-100.7-4.4%
weak2,29849.2%47.1% – 51.2%51.7%+2.5 pts-175.2-7.6%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
6.4%-38.5%86234.6%46.1%-11.5 pts
38.5%-42.1%86240.5%46.5%-6.1 pts
42.1%-44.7%86243.5%46.6%-3.2 pts
44.7%-47.0%86145.9%49.7%-3.8 pts
47.0%-49.3%86248.1%50.8%-2.7 pts
49.3%-52.0%86250.6%49.2%+1.4 pts
52.0%-55.3%86253.5%48.4%+5.1 pts
55.3%-80.1%86258.8%53.5%+5.3 pts
NRFI* UNGRADED — no price history was ever collected, so there is nothing to grade against
Winning? Cannot be answered. No historical price exists for this market, so there is no break-even to measure against — win rate is 51.5% over 7,201 picks, and a perfectly calibrated 52% laid at -125 loses all season.
Calibrated? Broadly overconfident — it claimed 53.8% and delivered 51.5%, 2.2 points too high. Worst on its strong picks: claimed 57.1%, got 52.4%.
Bands separating? No — they do not even order: weak is beating mid. A strength label that runs backwards is carrying no information.
Big enough to say? Yes — 7,201 picks, interval 2.3 points wide (50.4% to 52.7%).
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff by
strong2,40152.4%50.4% – 54.4%57.1%+4.7 pts
mid2,40050.9%48.9% – 52.9%53.2%+2.3 pts
weak2,40051.3%49.3% – 53.3%51.0%-0.3 pts
▸ Show 2026 rebuilt after the fact — an upper bound, never a record

2026, rebuilt afterwards. These picks were never published — an upper bound, never evidence.

Moneyline validated — but HIGH and MEDIUM no longer separate
Winning? No — -99.3 units over 1,442 priced picks, a -6.9% return at the prices actually paid. Win rate 52.5%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly overconfident — it claimed 56.5% and delivered 52.5%, 4.1 points too high. Worst on its HIGH picks: claimed 61.5%, got 52.8%.
Bands separating? No — they do not even order: MEDIUM is beating HIGH. A strength label that runs backwards is carrying no information.
Big enough to say? Yes — 1,445 picks, interval 5.1 points wide (49.9% to 55.0%).
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
HIGH26752.8%46.8% – 58.7%61.5%+8.7 pts-33.4-12.5%
MEDIUM49459.3%54.9% – 63.6%58.4%-1.0 pts+15.43.1%
LOW68447.4%43.7% – 51.1%53.3%+5.9 pts-81.3-11.9%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
14.7%-46.0%18141.3%48.6%-7.3 pts
46.0%-49.7%18048.0%51.7%-3.6 pts
49.7%-52.1%18151.0%43.6%+7.4 pts
52.1%-54.3%18053.2%47.2%+6.0 pts
54.3%-56.2%18155.3%48.6%+6.6 pts
56.3%-58.5%18057.3%57.8%-0.5 pts
58.6%-61.6%18160.0%56.9%+3.1 pts
61.6%-77.3%18164.7%63.0%+1.7 pts
Total* TESTED AND FAILED — bands measured on 8,272 games and did not separate
Winning? No — -41.4 units over 1,375 priced picks, a -3.0% return at the prices actually paid. Win rate 51.1%, which is only meaningful against the break-even, never against 50%.
Calibrated? Broadly overconfident — it claimed 56.1% and delivered 51.1%, 5.0 points too high. Worst on its strong picks: claimed 61.5%, got 48.6%.
Bands separating? No — they do not even order: mid is beating strong. A strength label that runs backwards is carrying no information.
Big enough to say? Yes — 1,377 picks, interval 5.3 points wide (48.4% to 53.7%).
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
strong45948.6%44.0% – 53.1%61.5%+12.9 pts-36.2-7.9%
mid45952.9%48.4% – 57.5%55.0%+2.1 pts+1.40.3%
weak45951.6%47.1% – 56.2%51.6%-0.0 pts-6.6-1.4%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
8.3%-39.5%17234.4%52.3%-17.9 pts
39.5%-42.2%17240.9%51.7%-10.8 pts
42.2%-44.6%17243.5%45.3%-1.8 pts
44.6%-46.6%17245.6%51.2%-5.5 pts
46.7%-48.8%17247.7%50.0%-2.3 pts
48.8%-51.3%17250.1%50.6%-0.5 pts
51.3%-54.1%17252.6%51.7%+0.9 pts
54.1%-67.7%17357.4%52.6%+4.8 pts
NRFI* UNGRADED — no price history was ever collected, so there is nothing to grade against
Winning? Cannot be answered. No historical price exists for this market, so there is no break-even to measure against — win rate is 53.1% over 1,445 picks, and a perfectly calibrated 52% laid at -125 loses all season.
Calibrated? Broadly honest — it claimed 54.7% and delivered 53.1%, 1.6 points too high. Worst on its strong picks: claimed 59.1%, got 55.2%.
Bands separating? They ORDER but do not SEPARATE — strong 55.2% down to weak 51.2%, but the intervals for strong and mid, mid and weak overlap, so those are not distinguishable.
Big enough to say? Yes — 1,445 picks, interval 5.1 points wide (50.6% to 55.7%).
▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff by
strong48255.2%50.7% – 59.6%59.1%+3.9 pts
mid48153.0%48.5% – 57.4%54.0%+1.0 pts
weak48251.2%46.8% – 55.7%51.2%-0.1 pts
▸ Show the full technical report (fixed-width, as the terminal prints it)
====================================================================================================
TRACKING — THREE SEPARATE LEDGERS
====================================================================================================
These three markets are NEVER pooled. Different base rates, different
validation status, different meanings. There is no combined record
anywhere in this report, by design.

Live begins 2026-08-10 — the first slate a card actually published, read from the ledger. Backtest and live are never summed.

####################################################################################################
# BACKTEST
# 2021-2025 walk-forward. The model never saw these rows, but the
# feature pipeline was designed while looking at this era.
####################################################################################################

----------------------------------------------------------------------------------------------------
MONEYLINE — the only validated market. The band claim below is
            measured on this run, not quoted from memory.
            Measured on 8,650 games: HIGH 59.4% / MEDIUM 57.1% / LOW 52.2%.
            They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM overlap.
            Clean: MEDIUM over LOW.
            Intervals: HIGH [57.1%,61.8%] n=1,657, MEDIUM [55.2%,58.9%] n=2,772, LOW [50.7%,53.7%] n=4,221.
----------------------------------------------------------------------------------------------------
  hit rate by band, with 95% CI, stated vs actual, and units.
  n = all graded games (hit rate, calibration).
  n_priced = the subset with a real price (units, ROI). The 2026
  archive backfill priced 2,299 of the 2,301 games that had none,
  so n_priced now nearly equals n; the two games still missing
  start before the snapshot that was bought.
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual   units   roi  n_priced
      HIGH 1390    60.7% 58.1% 63.3%     5.1%  62.4%            +1.7 pts  -54.16 -3.9%      1389
    MEDIUM 2278    56.6% 54.5% 58.6%     4.1%  58.0%            +1.4 pts  -96.22 -4.2%      2276
       LOW 3537    53.2% 51.5% 54.8%     3.3%  53.7%            +0.5 pts -143.44 -4.1%      3527

  calibration by probability bucket:
         bucket   n stated actual      gap ci_lo ci_hi
    13.6%-44.2% 901  39.5%  44.2% -4.6 pts 41.0% 47.4%
    44.2%-48.1% 900  46.3%  46.1% +0.2 pts 42.9% 49.4%
    48.1%-50.8% 901  49.5%  52.6% -3.1 pts 49.3% 55.9%
    50.8%-53.3% 900  52.1%  51.0% +1.1 pts 47.7% 54.3%
    53.3%-55.6% 901  54.4%  54.4% +0.0 pts 51.1% 57.6%
    55.6%-58.2% 900  56.8%  56.7% +0.2 pts 53.4% 59.9%
    58.2%-61.8% 901  59.8%  58.9% +0.9 pts 55.7% 62.1%
    61.8%-80.9% 901  65.8%  62.5% +3.3 pts 59.3% 65.6%

----------------------------------------------------------------------------------------------------
TOTALS* — BANDS WERE TESTED AND FAILED. Incremental R^2 over the
          posted line is +0.0022, and weak/mid/strong is CLAIM
          STRENGTH ONLY, never a validated band.
          Measured on 8,272 games: HIGH 52.5% / MEDIUM 52.5% / LOW 50.1%.
          They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM, MEDIUM and LOW overlap.
          Intervals: HIGH [50.0%,55.1%] n=1,513, MEDIUM [50.7%,54.4%] n=2,754, LOW [48.6%,51.7%] n=4,005.
----------------------------------------------------------------------------------------------------
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual   units   roi  n_priced
    strong 2299    54.1% 52.1% 56.1%     4.1%  61.9%            +7.8 pts   25.90  1.1%      2299
       mid 2298    51.0% 49.0% 53.0%     4.1%  55.3%            +4.3 pts -100.74 -4.4%      2298
      weak 2298    49.2% 47.1% 51.2%     4.1%  51.7%            +2.5 pts -175.21 -7.6%      2298

  calibration by probability bucket:
         bucket   n stated actual       gap ci_lo ci_hi
     6.4%-38.5% 862  34.6%  46.1% -11.5 pts 42.8% 49.4%
    38.5%-42.1% 862  40.5%  46.5%  -6.1 pts 43.2% 49.9%
    42.1%-44.7% 862  43.5%  46.6%  -3.2 pts 43.3% 50.0%
    44.7%-47.0% 861  45.9%  49.7%  -3.8 pts 46.4% 53.0%
    47.0%-49.3% 862  48.1%  50.8%  -2.7 pts 47.5% 54.1%
    49.3%-52.0% 862  50.6%  49.2%  +1.4 pts 45.9% 52.5%
    52.0%-55.3% 862  53.5%  48.4%  +5.1 pts 45.1% 51.7%
    55.3%-80.1% 862  58.8%  53.5%  +5.3 pts 50.1% 56.8%

----------------------------------------------------------------------------------------------------
NRFI* — UNGRADED. No historical price is HELD for this market, so
        break-even, units and market comparison are UNCOMPUTABLE,
        not merely absent. Hit rate only.
        (A season of history is purchasable at ~17,210 credits and
         has not been bought. The LIVE market is obtainable:
         DraftKings quoted the 0.5 line on 3 of the 3 events
         carrying it on 2026-08-13, under the market key
         alternate_totals_1st_1_innings.)
----------------------------------------------------------------------------------------------------
     group    n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual
    strong 2401    52.4% 50.4% 54.4%     4.0%  57.1%            +4.7 pts
       mid 2400    50.9% 48.9% 52.9%     4.0%  53.2%            +2.3 pts
      weak 2400    51.3% 49.3% 53.3%     4.0%  51.0%            -0.3 pts

####################################################################################################
# RECONSTRUCTED
# 2026 REBUILT FROM WALK-FORWARD OUTPUT. NOT a live record: these
# picks were never published, the model was re-run afterwards on a
# different fold, and the pipeline was designed with these results
# already visible. An upper bound, never evidence.
# The live record lives in data/live/published_picks.csv and is shown on the card.
####################################################################################################

----------------------------------------------------------------------------------------------------
MONEYLINE — the only validated market. The band claim below is
            measured on this run, not quoted from memory.
            Measured on 8,650 games: HIGH 59.4% / MEDIUM 57.1% / LOW 52.2%.
            They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM overlap.
            Clean: MEDIUM over LOW.
            Intervals: HIGH [57.1%,61.8%] n=1,657, MEDIUM [55.2%,58.9%] n=2,772, LOW [50.7%,53.7%] n=4,221.
----------------------------------------------------------------------------------------------------
  hit rate by band, with 95% CI, stated vs actual, and units.
  n = all graded games (hit rate, calibration).
  n_priced = the subset with a real price (units, ROI). The 2026
  archive backfill priced 2,299 of the 2,301 games that had none,
  so n_priced now nearly equals n; the two games still missing
  start before the snapshot that was bought.
     group   n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual  units    roi  n_priced
      HIGH 267    52.8% 46.8% 58.7%    11.9%  61.5%            +8.7 pts -33.42 -12.5%       267
    MEDIUM 494    59.3% 54.9% 63.6%     8.6%  58.4%            -1.0 pts  15.38   3.1%       491
       LOW 684    47.4% 43.7% 51.1%     7.5%  53.3%            +5.9 pts -81.30 -11.9%       684
    HIGH: !  n=267 — interval is wide; treat as provisional

  calibration by probability bucket:
         bucket   n stated actual      gap ci_lo ci_hi
    14.7%-46.0% 181  41.3%  48.6% -7.3 pts 41.4% 55.9%
    46.0%-49.7% 180  48.0%  51.7% -3.6 pts 44.4% 58.9%
    49.7%-52.1% 181  51.0%  43.6% +7.4 pts 36.6% 50.9%
    52.1%-54.3% 180  53.2%  47.2% +6.0 pts 40.1% 54.5%
    54.3%-56.2% 181  55.3%  48.6% +6.6 pts 41.4% 55.9%
    56.3%-58.5% 180  57.3%  57.8% -0.5 pts 50.5% 64.8%
    58.6%-61.6% 181  60.0%  56.9% +3.1 pts 49.6% 63.9%
    61.6%-77.3% 181  64.7%  63.0% +1.7 pts 55.7% 69.7%

----------------------------------------------------------------------------------------------------
TOTALS* — BANDS WERE TESTED AND FAILED. Incremental R^2 over the
          posted line is +0.0022, and weak/mid/strong is CLAIM
          STRENGTH ONLY, never a validated band.
          Measured on 8,272 games: HIGH 52.5% / MEDIUM 52.5% / LOW 50.1%.
          They ORDER but do not SEPARATE: the intervals for HIGH and MEDIUM, MEDIUM and LOW overlap.
          Intervals: HIGH [50.0%,55.1%] n=1,513, MEDIUM [50.7%,54.4%] n=2,754, LOW [48.6%,51.7%] n=4,005.
----------------------------------------------------------------------------------------------------
     group   n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual  units   roi  n_priced
    strong 459    48.6% 44.0% 53.1%     9.1%  61.5%           +12.9 pts -36.18 -7.9%       459
       mid 459    52.9% 48.4% 57.5%     9.1%  55.0%            +2.1 pts   1.35  0.3%       457
      weak 459    51.6% 47.1% 56.2%     9.1%  51.6%            -0.0 pts  -6.61 -1.4%       459

  calibration by probability bucket:
         bucket   n stated actual       gap ci_lo ci_hi
     8.3%-39.5% 172  34.4%  52.3% -17.9 pts 44.9% 59.7%
    39.5%-42.2% 172  40.9%  51.7% -10.8 pts 44.3% 59.1%
    42.2%-44.6% 172  43.5%  45.3%  -1.8 pts 38.1% 52.8%
    44.6%-46.6% 172  45.6%  51.2%  -5.5 pts 43.7% 58.5%
    46.7%-48.8% 172  47.7%  50.0%  -2.3 pts 42.6% 57.4%
    48.8%-51.3% 172  50.1%  50.6%  -0.5 pts 43.2% 58.0%
    51.3%-54.1% 172  52.6%  51.7%  +0.9 pts 44.3% 59.1%
    54.1%-67.7% 173  57.4%  52.6%  +4.8 pts 45.2% 59.9%

----------------------------------------------------------------------------------------------------
NRFI* — UNGRADED. No historical price is HELD for this market, so
        break-even, units and market comparison are UNCOMPUTABLE,
        not merely absent. Hit rate only.
        (A season of history is purchasable at ~17,210 credits and
         has not been bought. The LIVE market is obtainable:
         DraftKings quoted the 0.5 line on 3 of the 3 events
         carrying it on 2026-08-13, under the market key
         alternate_totals_1st_1_innings.)
----------------------------------------------------------------------------------------------------
     group   n hit_rate ci_lo ci_hi ci_width stated stated_minus_actual
    strong 482    55.2% 50.7% 59.6%     8.8%  59.1%            +3.9 pts
       mid 481    53.0% 48.5% 57.4%     8.9%  54.0%            +1.0 pts
      weak 482    51.2% 46.8% 55.7%     8.9%  51.2%            -0.1 pts

====================================================================================================
READING THESE NUMBERS
====================================================================================================
  Every hit rate above carries a 95% interval. A rate whose interval
  contains the relevant break-even (or 0.50) has not demonstrated
  anything, however good the point estimate looks.
  Live moneyline has 39 settled picks over 3 slates (about 13 a slate), and at that count the 95% interval is +/- 15.0 points.
  Narrowing it to +/- 3 points — the width at which a live rate could be told apart from the no-model base rate — needs roughly 1,066 settled picks, about 79 more slates at the current rate.
  Live conclusions should not be drawn before then.