MLB card — 2026-08-13

9 games · times in Phoenix · generated 2026-08-13 15:07 UTC
Side is the only validated column. Total* and NRFI* are not — see the notes at the bottom.
Probabilities: calibrated on 2,074 games from the prior year.

Season to datethree markets, never pooled · LIVE since 2026-08-10

2026 backtest at real prices: -6.1% ROI over 1,443 games. The loss was statistically indistinguishable from paying the vig with zero skill (residual -2.1%, z=-0.90).

MONEYLINE
 nrecordhit %95% CIunitsb/e
all4029-1172.5%57.2 – 83.9+13.954.4%
HIGH32-1n<30+0.162.1%
MEDIUM1210-2n<30+4.353.7%
LOW2517-8n<30+9.553.4%
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other. RECONSTRUCTED is 2026 rebuilt after the fact — an upper bound, not a record. Backtest is the era the features were designed against.

 nrecordhit %95% CIunitsb/e
RECONSTRUCTED · all1445758-68752.5%49.9 – 55.0-99.356.5%
backtest · all72054014-319155.7%54.6 – 56.9-293.858.3%
TOTAL* Checked and it did not work.
▸ Show the measurements

Bands were measured on 8,272 out-of-sample games and did not separate. HIGH hit 0.5223 with a 95% CI of [0.4972, 0.5473] — an interval that contains 0.50 and overlaps LOW. The bands do not even order: MEDIUM 0.5307 sits above HIGH 0.5223. Incremental R-squared over the posted line is +0.0022. Calibrated but uninformative.

 nrecordhit %95% CIunitsb/e
all00-0no graded picks
strong00-0no graded picks
mid00-0no graded picks
weak00-0no graded picks
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other. RECONSTRUCTED is 2026 rebuilt after the fact — an upper bound, not a record. Backtest is the era the features were designed against.

 nrecordhit %95% CIunitsb/e
RECONSTRUCTED · all1377703-67451.1%48.4 – 53.7-41.452.5%
backtest · all68953546-334951.4%50.2 – 52.6-250.053.4%

40 totals picks (2026-08-10 to 2026-08-12) can never be settled: they were published before the ledger recorded the posted line, and that line is not recoverable after the fact. They are not back-filled. Totals published from now on carry their line and settle normally.

NRFI* Never checked — there is no price to check it against.
▸ Show the measurements

No historical NRFI price exists anywhere, so hit-rate-against-break-even and Brier-against-market are uncomputable, not merely poor. Two of the three metric tiers cannot run, so there is no break-even column and no band.

 nrecordhit %95% CI
all4026-1465.0%49.5 – 77.9
strong138-5n<30
mid1310-3n<30
weak148-6n<30
▸ Show backtest and reconstructed (context, not live performance)

Never added to the live rows above, and never to each other. RECONSTRUCTED is 2026 rebuilt after the fact — an upper bound, not a record. Backtest is the era the features were designed against.

 nrecordhit %95% CI
RECONSTRUCTED · all1445768-67753.1%50.6 – 55.7
backtest · all72013712-348951.5%50.4 – 52.7

These three are never added together. Only LIVE is evidence.

▸ Show what these three phases mean

LIVE — Read from data/live/published_picks.csv, which is written at publish time and refuses any pick for a game that has already started. It is the only source that can support a forward claim.

RECONSTRUCTED — 2026 walk-forward output. Those picks came from a model fit on a different fold than the card uses, and every design choice in the pipeline was made with these results already visible. Treat as an upper bound, never as evidence.

backtest — 2021-2025 walk-forward. The model never saw these rows, but the feature pipeline was built while looking at this era.

What to trust

Side is the only column checked against real results — and it came out in the right order.

Total and NRFI have not been checked. The model has an opinion; nobody has verified it.

What the confidence words mean

HIGH / MED / LOW on Side were measured against what actually happened.

weak / mid / strong on Total and NRFI are only how sure the model feels. Untested.

Different words on purpose, so one can never be read as the other.

The honest headline

Tested on 1,443 real 2026 games at real prices, this lost 6.1%.

That is what paying the bookmaker’s cut (the “vig”) costs you, with no skill at all.

Nothing here has been shown to make money.

▸ Full technical detail

TOTAL* — tested and FAILED. Bands were measured on 8,272 out-of-sample games and did not separate. HIGH hit 0.5223 with a 95% CI of [0.4972, 0.5473] — an interval that contains 0.50 and overlaps LOW. Incremental R² over the posted line is +0.0022. Calibrated but uninformative. The 2026 backfill added ~2,200 priced games and the verdict did not move.

NRFI* — UNGRADED, never tested at all. No historical NRFI price exists anywhere, so hit-rate-against-break-even and Brier-against-market are uncomputable, not merely poor. No break-even is shown because there is no price.

weak / mid / strong is model conviction, NOT validation. It is which third of the claim-strength distribution the number falls in, and has not been shown to predict anything. A “strong” total is not a HIGH side.

SIDE re-measured on 8,650 games: HIGH 0.6001 [0.5760, 0.6243], MEDIUM 0.5693 [0.5505, 0.5881], LOW 0.5197 [0.5049, 0.5345]. The ordering holds, but HIGH and MEDIUM now overlap — on the earlier 6,367-game window they did not. HIGH beats MEDIUM on the point estimate and is not proven distinct from it; MEDIUM over LOW is still clean. Only side HIGH is shaded. Break-even is shown beside every priced pick and never filters — every game gets an output, there is no “no bet”.

When Total* is empty there are two reasons. Either no posted total exists, in which case the model has nothing to read P(over) off and the blank is a missing input; or a line exists but the game is inside the 21-day sealed window that is deliberately not graded. Since the archive backfill posted lines for 2,299 games of the 2026 season, the second is the usual reason on the newest slate.

Probabilities are recalibrated against the trailing year. The raw model runs overconfident on recent data — reliability 0.0006 on 2022–2025 against 0.0043 on 2026, always the same direction. The correction is a Platt map fitted strictly earlier than the slate it is applied to, and it is monotone, so band ordering is untouched: only the numbers move. On 2026 it cuts reliability to 0.0026 and closes the HIGH band’s stated-minus-actual gap from +0.0298 to +0.0012. It does not fully fix the drift.

How the model has actually done

Three markets, never pooled — different base rates, different validation status, different meanings. There is deliberately no combined number anywhere below. Backtest and live are never added together either.

Backtest

2021–2025. The model never saw these games, but the features were designed while looking at this era.

Moneyline validated — but HIGH and MEDIUM no longer separate

Strongest picks won 60.7% of the time, weakest 53.2% (a +7.5 point spread). The model was most overconfident on its HIGH picks, claiming 62.4% where 60.7% happened. Across every strength this period, -293.8 units.

▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
HIGH1,39060.7%58.1% – 63.3%62.4%+1.7 pts-54.2-3.9%
MEDIUM2,27856.6%54.5% – 58.6%58.0%+1.4 pts-96.2-4.2%
LOW3,53753.2%51.5% – 54.8%53.7%+0.5 pts-143.4-4.1%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
0.136-0.44290139.5%44.2%-4.6 pts
0.442-0.48190046.3%46.1%+0.2 pts
0.481-0.50890149.5%52.6%-3.1 pts
0.508-0.53390052.1%51.0%+1.1 pts
0.533-0.55690154.4%54.4%+0.0 pts
0.556-0.58290056.8%56.7%+0.2 pts
0.582-0.61890159.8%58.9%+0.9 pts
0.618-0.80990165.8%62.5%+3.3 pts
Total* TESTED AND FAILED — bands measured on 8,272 games and did not separate

Strongest picks won 54.1% of the time, weakest 49.2% (a +4.9 point spread). The model was most overconfident on its strong picks, claiming 61.9% where 54.1% happened. Across every strength this period, -250.0 units.

▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
strong2,29954.1%52.1% – 56.1%61.9%+7.8 pts+25.91.1%
mid2,29851.0%49.0% – 53.0%55.3%+4.3 pts-100.7-4.4%
weak2,29849.2%47.1% – 51.2%51.7%+2.5 pts-175.2-7.6%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
0.064-0.38586234.6%46.1%-11.5 pts
0.385-0.42186240.5%46.5%-6.1 pts
0.421-0.44786243.5%46.6%-3.2 pts
0.447-0.47086145.9%49.7%-3.8 pts
0.470-0.49386248.1%50.8%-2.7 pts
0.493-0.52086250.6%49.2%+1.4 pts
0.520-0.55386253.5%48.4%+5.1 pts
0.553-0.80186258.8%53.5%+5.3 pts
NRFI* UNGRADED — no price exists anywhere, so this has never been tested against a market

Strongest picks won 52.4% of the time, weakest 51.3% (a +1.1 point spread). The model was most overconfident on its strong picks, claiming 57.1% where 52.4% happened.

▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff by
strong2,40152.4%50.4% – 54.4%57.1%+4.7 pts
mid2,40050.9%48.9% – 52.9%53.2%+2.3 pts
weak2,40051.3%49.3% – 53.3%51.0%-0.3 pts
Reconstructed

2026, rebuilt afterwards. These picks were never published — an upper bound, never evidence.

Moneyline validated — but HIGH and MEDIUM no longer separate

Strongest picks won 52.8% of the time, weakest 47.4% (a +5.4 point spread). The model was most overconfident on its HIGH picks, claiming 61.5% where 52.8% happened. Across every strength this period, -99.3 units.

▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
HIGH26752.8%46.8% – 58.7%61.5%+8.7 pts-33.4-12.5%
MEDIUM49459.3%54.9% – 63.6%58.4%-1.0 pts+15.43.1%
LOW68447.4%43.7% – 51.1%53.3%+5.9 pts-81.3-11.9%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
0.147-0.46018141.3%48.6%-7.3 pts
0.460-0.49718048.0%51.7%-3.6 pts
0.497-0.52118151.0%43.6%+7.4 pts
0.521-0.54318053.2%47.2%+6.0 pts
0.543-0.56218155.3%48.6%+6.6 pts
0.563-0.58518057.3%57.8%-0.5 pts
0.586-0.61618160.0%56.9%+3.1 pts
0.616-0.77318164.7%63.0%+1.7 pts
Total* TESTED AND FAILED — bands measured on 8,272 games and did not separate

Strongest picks won 48.6% of the time, weakest 51.6% (a -3.1 point spread). The model was most overconfident on its strong picks, claiming 61.5% where 48.6% happened. Across every strength this period, -41.4 units.

▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff byunitsreturn
strong45948.6%44.0% – 53.1%61.5%+12.9 pts-36.2-7.9%
mid45952.9%48.4% – 57.5%55.0%+2.1 pts+1.40.3%
weak45951.6%47.1% – 56.2%51.6%-0.0 pts-6.6-1.4%
▸ Show whether the stated probabilities were honest

Each row is a group of games the model felt similarly about. If it is calibrated, the two middle columns match.

when the model said…gamesit claimed…this happenedoff by
0.083-0.39517234.4%52.3%-17.9 pts
0.395-0.42217240.9%51.7%-10.8 pts
0.422-0.44617243.5%45.3%-1.8 pts
0.446-0.46617245.6%51.2%-5.5 pts
0.467-0.48817247.7%50.0%-2.3 pts
0.488-0.51317250.1%50.6%-0.5 pts
0.513-0.54117252.6%51.7%+0.9 pts
0.541-0.67717357.4%52.6%+4.8 pts
NRFI* UNGRADED — no price exists anywhere, so this has never been tested against a market

Strongest picks won 55.2% of the time, weakest 51.2% (a +3.9 point spread). The model was most overconfident on its strong picks, claiming 59.1% where 55.2% happened.

▸ Show the numbers
strengthgameshit rate95% CImodel claimedoff by
strong48255.2%50.7% – 59.6%59.1%+3.9 pts
mid48153.0%48.5% – 57.4%54.0%+1.0 pts
weak48251.2%46.8% – 55.7%51.2%-0.1 pts
▸ Show the full technical report (fixed-width, as the terminal prints it)
====================================================================================================
TRACKING — THREE SEPARATE LEDGERS
====================================================================================================
These three markets are NEVER pooled. Different base rates, different
validation status, different meanings. There is no combined record
anywhere in this report, by design.

Live begins 2026-01-01 (first scheduled 15:30 UTC run). Backtest and live are never summed.

####################################################################################################
# BACKTEST
# 2021-2025 walk-forward. The model never saw these rows, but the
# feature pipeline was designed while looking at this era.
####################################################################################################

----------------------------------------------------------------------------------------------------
MONEYLINE — VALIDATED, but the claim has weakened. Re-measured on
            8,650 games: HIGH 0.6001 / MEDIUM 0.5693 / LOW 0.5197.
            Ordering holds; HIGH and MEDIUM intervals now OVERLAP
            ([0.5760,0.6243] vs [0.5505,0.5881]), where the earlier
            6,367-game window showed none. MEDIUM > LOW is clean.
----------------------------------------------------------------------------------------------------
  hit rate by band, with 95% CI, stated vs actual, and units.
  n = all graded games (hit rate, calibration).
  n_priced = the subset with a real price (units, ROI). The 2026
  archive backfill priced 2,299 of the 2,301 games that had none,
  so n_priced now nearly equals n; the two games still missing
  start before the snapshot that was bought.
     group    n  hit_rate  ci_lo  ci_hi  ci_width  stated  stated_minus_actual   units     roi  n_priced
      HIGH 1390    0.6072 0.5813 0.6325    0.0513  0.6244               0.0172  -54.16 -0.0390      1389
    MEDIUM 2278    0.5658 0.5454 0.5861    0.0407  0.5795               0.0137  -96.22 -0.0423      2276
       LOW 3537    0.5318 0.5153 0.5482    0.0329  0.5370               0.0052 -143.44 -0.0407      3527

  calibration by probability bucket:
         bucket   n  stated   actual       gap  ci_lo  ci_hi
    0.136-0.442 901  0.3954 0.441731 -0.046333 0.4096 0.4743
    0.442-0.481 900  0.4627 0.461111  0.001566 0.4288 0.4938
    0.481-0.508 901  0.4951 0.526082 -0.031031 0.4934 0.5585
    0.508-0.533 900  0.5207 0.510000  0.010722 0.4774 0.5425
    0.533-0.556 901  0.5443 0.543840  0.000459 0.5112 0.5761
    0.556-0.582 900  0.5683 0.566667  0.001670 0.5341 0.5987
    0.582-0.618 901  0.5981 0.589345  0.008763 0.5569 0.6210
    0.618-0.809 901  0.6576 0.624861  0.032784 0.5928 0.6559

----------------------------------------------------------------------------------------------------
TOTALS* — BANDS WERE TESTED AND FAILED.
          HIGH hit 0.5223, 95% CI [0.4972, 0.5473] — an interval that
          CONTAINS 0.50 and overlaps LOW. Incremental R^2 over the
          posted line is +0.0022. weak/mid/strong below is CLAIM
          STRENGTH ONLY, not a validated band.
----------------------------------------------------------------------------------------------------
     group    n  hit_rate  ci_lo  ci_hi  ci_width  stated  stated_minus_actual   units     roi  n_priced
    strong 2299    0.5411 0.5207 0.5614    0.0407  0.6195               0.0784   25.90  0.0113      2299
       mid 2298    0.5100 0.4896 0.5304    0.0408  0.5533               0.0433 -100.74 -0.0438      2298
      weak 2298    0.4917 0.4713 0.5122    0.0408  0.5168               0.0251 -175.21 -0.0762      2298

  calibration by probability bucket:
         bucket   n  stated   actual       gap  ci_lo  ci_hi
    0.064-0.385 862  0.3456 0.460557 -0.114930 0.4275 0.4939
    0.385-0.421 862  0.4046 0.465197 -0.060630 0.4321 0.4986
    0.421-0.447 862  0.4346 0.466357 -0.031796 0.4333 0.4997
    0.447-0.470 861  0.4590 0.497096 -0.038090 0.4638 0.5304
    0.470-0.493 862  0.4814 0.508121 -0.026719 0.4748 0.5414
    0.493-0.520 862  0.5057 0.491879  0.013807 0.4586 0.5252
    0.520-0.553 862  0.5349 0.483759  0.051104 0.4505 0.5171
    0.553-0.801 862  0.5881 0.534803  0.053280 0.5014 0.5679

----------------------------------------------------------------------------------------------------
NRFI* — UNGRADED. No historical price exists for this market, so
        break-even, units and market comparison are UNCOMPUTABLE,
        not merely absent. Hit rate only.
----------------------------------------------------------------------------------------------------
     group    n  hit_rate  ci_lo  ci_hi  ci_width  stated  stated_minus_actual
    strong 2401    0.5244 0.5044 0.5443    0.0399  0.5709               0.0465
       mid 2400    0.5092 0.4892 0.5291    0.0400  0.5325               0.0233
      weak 2400    0.5129 0.4929 0.5329    0.0400  0.5100              -0.0029

####################################################################################################
# RECONSTRUCTED
# 2026 REBUILT FROM WALK-FORWARD OUTPUT. NOT a live record: these
# picks were never published, the model was re-run afterwards on a
# different fold, and the pipeline was designed with these results
# already visible. An upper bound, never evidence.
# The live record lives in data/live/published_picks.csv and is shown on the card.
####################################################################################################

----------------------------------------------------------------------------------------------------
MONEYLINE — VALIDATED, but the claim has weakened. Re-measured on
            8,650 games: HIGH 0.6001 / MEDIUM 0.5693 / LOW 0.5197.
            Ordering holds; HIGH and MEDIUM intervals now OVERLAP
            ([0.5760,0.6243] vs [0.5505,0.5881]), where the earlier
            6,367-game window showed none. MEDIUM > LOW is clean.
----------------------------------------------------------------------------------------------------
  hit rate by band, with 95% CI, stated vs actual, and units.
  n = all graded games (hit rate, calibration).
  n_priced = the subset with a real price (units, ROI). The 2026
  archive backfill priced 2,299 of the 2,301 games that had none,
  so n_priced now nearly equals n; the two games still missing
  start before the snapshot that was bought.
     group   n  hit_rate  ci_lo  ci_hi  ci_width  stated  stated_minus_actual  units     roi  n_priced
      HIGH 267    0.5281 0.4682 0.5871    0.1189  0.6148               0.0867 -33.42 -0.1252       267
    MEDIUM 494    0.5931 0.5492 0.6356    0.0863  0.5836              -0.0095  15.38  0.0313       491
       LOW 684    0.4737 0.4365 0.5111    0.0746  0.5325               0.0588 -81.30 -0.1189       684
    HIGH: !  n=267 — interval is wide; treat as provisional

  calibration by probability bucket:
         bucket   n  stated   actual       gap  ci_lo  ci_hi
    0.147-0.460 181  0.4134 0.486188 -0.072740 0.4144 0.5585
    0.460-0.497 180  0.4802 0.516667 -0.036468 0.4441 0.5886
    0.497-0.521 181  0.5101 0.436464  0.073625 0.3663 0.5093
    0.521-0.543 180  0.5321 0.472222  0.059837 0.4006 0.5450
    0.543-0.562 181  0.5526 0.486188  0.066431 0.4144 0.5585
    0.563-0.585 180  0.5730 0.577778 -0.004745 0.5047 0.6476
    0.586-0.616 181  0.6000 0.569061  0.030988 0.4962 0.6390
    0.616-0.773 181  0.6467 0.629834  0.016915 0.5575 0.6968

----------------------------------------------------------------------------------------------------
TOTALS* — BANDS WERE TESTED AND FAILED.
          HIGH hit 0.5223, 95% CI [0.4972, 0.5473] — an interval that
          CONTAINS 0.50 and overlaps LOW. Incremental R^2 over the
          posted line is +0.0022. weak/mid/strong below is CLAIM
          STRENGTH ONLY, not a validated band.
----------------------------------------------------------------------------------------------------
     group   n  hit_rate  ci_lo  ci_hi  ci_width  stated  stated_minus_actual  units     roi  n_priced
    strong 459    0.4858 0.4404 0.5315    0.0911  0.6153               0.1295 -36.18 -0.0788       459
       mid 459    0.5294 0.4837 0.5746    0.0909  0.5504               0.0210   1.35  0.0030       457
      weak 459    0.5163 0.4707 0.5617    0.0911  0.5161              -0.0002  -6.61 -0.0144       459

  calibration by probability bucket:
         bucket   n  stated   actual       gap  ci_lo  ci_hi
    0.083-0.395 172  0.3438 0.523256 -0.179452 0.4489 0.5966
    0.395-0.422 172  0.4090 0.517442 -0.108450 0.4432 0.5909
    0.422-0.446 172  0.4352 0.453488 -0.018243 0.3809 0.5281
    0.446-0.466 172  0.4564 0.511628 -0.055228 0.4375 0.5853
    0.467-0.488 172  0.4772 0.500000 -0.022793 0.4261 0.5739
    0.488-0.513 172  0.5011 0.505814 -0.004741 0.4318 0.5796
    0.513-0.541 172  0.5260 0.517442  0.008536 0.4432 0.5909
    0.541-0.677 173  0.5742 0.526012  0.048173 0.4518 0.5990

----------------------------------------------------------------------------------------------------
NRFI* — UNGRADED. No historical price exists for this market, so
        break-even, units and market comparison are UNCOMPUTABLE,
        not merely absent. Hit rate only.
----------------------------------------------------------------------------------------------------
     group   n  hit_rate  ci_lo  ci_hi  ci_width  stated  stated_minus_actual
    strong 482    0.5519 0.5072 0.5957    0.0884  0.5908               0.0389
       mid 481    0.5301 0.4855 0.5743    0.0889  0.5396               0.0095
      weak 482    0.5124 0.4679 0.5568    0.0889  0.5118              -0.0007

====================================================================================================
READING THESE NUMBERS
====================================================================================================
  Every hit rate above carries a 95% interval. A rate whose interval
  contains the relevant break-even (or 0.50) has not demonstrated
  anything, however good the point estimate looks.
  At ~140 live picks by season end the interval is about +/- 8 points,
  which is wider than any edge worth having. Live conclusions should
  not be drawn this season.