How each matchup alert has done once the game finished, checked against the final score and the sportsbook number.
Games with no posted sportsbook number are left out of the win-loss counts.
What the model total has meant in tracked games.
| We said | Call | Record |
|---|---|---|
| 38 or less | Under 41Strongest | 8–0 (100%) |
| 44 or less | Under 45 | 14–1 (93%) |
| 45 to 54 | Over 45 | 6–1 (86%) |
| 55+ | Skip | 5–4 |
The first two rows overlap. Every game at 38 or less is also in the 44-or-less row. When both apply, Under 41 is the stronger play.
Loading records…
The EXPECTED chips on every game card, graded from the number we froze before kickoff.
The four stat badges, graded only when they were frozen before kickoff.
Situations that finished one way more often than games on similar closing spreads in our walk-forward backtest: our model's pregame numbers for 2020–25 (each built only from earlier games) against nflverse closing lines. Picked on 2020–24, checked on 2025.
A broad scan of 12,288 situation × result combinations turned up no more "winners" than shuffled results do, so only patterns with a football reason that also held in 2025 are kept. A 👀 Watching pattern becomes a ⚡ Pattern alert only after 25+ live games (after it was written down) with a 90% lower bound above the base by 3 points, and drops back when that stops being true.
Records only, not advice. 21+ only where legal.
Loading patterns…
Team-stat matchups (each team's season-to-date numbers from before kickoff, blended with last season in weeks 1–4) checked against nflverse results and closing lines, 2016–2025. Each rule was picked on 2016–22 only, then had to hold in 2023–24 and in the 2025 holdout, with no more than 2 losing seasons of 10. "Base" is the same result's rate in games on similar closing lines, so a lift is on top of what the market already priced (run- and pass-heavy games have no market; their base is the league rate on similar lines).
Multiple comparisons: we tried 40,566 rules (407 features × thresholds on round percentiles, plus two-condition combos) across 15 results. With that many tries, dozens pass by luck, so the whole search was re-run on shuffled results (shuffled within season and line band). Only results whose real rules clearly beat the shuffles keep any rule; the rest are listed below as "nothing beat chance". 🔥 Super trend = 2025 holdout lift of 8+ points, 2016–24 interval above the base and 7+ winning seasons. A trend is promoted live after 25+ games with a 90% lower bound 3 points above its base, and drops back (or off the cards) when its live record trails the base.
Live records start Sep 27; this season's earlier weeks are graded as backtest. Analysis and entertainment, not betting advice. 21+ only where legal.
When: The favorite punts at least 0.56 more times per game than the underdog (top 20% of games).
Why it could be real: The spread leans on the favorite's record and reputation, but its offense stalls more often than the underdog's; drive-by-drive production points to a closer game than the number.
2016–24 combined: 90% interval 55%–62% vs base 52%
Met the super-trend bar but kept at 👀 Watching: the model cross-check found no underdog-cover signal beyond the line, so only its live record can promote it.
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 10 of 10 seasons beat the base.
When: The underdog tries 0.8+ more fourth downs per game than the favorite (top 10%).
Why it could be real: Aggressive underdog coaches turn punts into extra possessions and points, and cautious favorites leave margin on the table, so these games tend to stay inside the number.
2016–24 combined: 90% interval 53%–63% vs base 50%
Met the super-trend bar but kept at 👀 Watching: the model cross-check found no underdog-cover signal beyond the line, so only its live record can promote it.
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 8 of 10 seasons beat the base.
When: The underdog tries 0.8+ more fourth downs per game than the favorite; and the two offenses average 22.25 or fewer drives per game combined (bottom third).
Why it could be real: Fewer possessions leave the better team fewer chances to pull away, and the aggressive underdog squeezes more out of each one: a shortened game helps the side getting points.
2016–24 combined: 90% interval 58%–71% vs base 50% · each condition alone (2016–22): 61% vs 52%, 57% vs 53%
Met the super-trend bar but kept at 👀 Watching: the model cross-check found no underdog-cover signal beyond the line, so only its live record can promote it.
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 8 of 10 seasons beat the base.
When: Both offenses project an explosive-play rate of about 10%+ against the other defense (top 20%).
Why it could be real: Explosive plays make results swingy: one long touchdown can keep an underdog in the game or produce a late back-door score, which favors the side getting points.
2016–24 combined: 90% interval 56%–63% vs base 52%
Kept at 👀 Watching: the model cross-check found no underdog-cover signal beyond the line, so only its live record can promote it.
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: Each team's rushing yards per game plus what the other defense allows add up to 252+ projected rushing yards (top 10%).
Why it could be real: When both rushing offenses meet run defenses that give up yards, both teams keep running; the betting total prices points, not how they are scored.
2016–24 combined: 90% interval 42%–51% vs base 27%
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: One team's rushing average and the opponent's rushing yards allowed average out to 140+ (top 10%).
Why it could be real: A strong ground game against a leaky run defense usually gets fed, and a team that is running well leans on it even more as it leads.
2016–24 combined: 90% interval 42%–52% vs base 27%
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: Both sides project 26.9+ carries (top 10%); and combined opponent-adjusted rushing efficiency in the top third.
Why it could be real: Two play-callers who run a lot, meeting defenses that can't stop it efficiently, pile up carries on both sides of the ball.
2016–24 combined: 90% interval 42%–56% vs base 25% · each condition alone (2016–22): 35% vs 25%, 42% vs 26%
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 8 of 10 seasons beat the base.
When: The underdog's yards per carry and the favorite's yards per carry allowed average 3.75 or less (bottom 10%).
Why it could be real: An underdog that can't run against this defense falls behind and has to throw, and the favorite answers through the air too, so passing yards pile up.
2016–24 combined: 90% interval 35%–45% vs base 30%
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 10 of 10 seasons beat the base.
When: Both sides project 24.1 or fewer carries (bottom third); and both sides project 3.8 yards per carry or less (bottom 20%).
Why it could be real: When neither offense runs often and neither can run efficiently in this matchup, both games plans go through the quarterback.
2016–24 combined: 90% interval 38%–49% vs base 31% · each condition alone (2016–22): 43% vs 37%, 38% vs 32%
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: Both sides project a 54%+ pass rate on early downs in neutral game states (top 10%).
Why it could be real: Early-down pass rate is a play-caller's habit; when both sides throw first even in neutral spots, the yardage goes through the air.
2016–24 combined: 90% interval 44%–55% vs base 39%
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 10 of 10 seasons beat the base.
Rules passing every check on the real results vs the average on shuffled results, and a cross-check model (regularised logistic regression, season-by-season cross-validation) with and without the team stats, scored on the 2025 holdout (AUC: 0.5 = no skill).
| Result | Passed | Shuffled avg | p | 2025 AUC line → +stats | Verdict |
|---|---|---|---|---|---|
| Low-scoring (≤37 and under the total) | 6 | 9.7 | 0.677 | 0.66 → 0.62 | Nothing beat chance |
| Under the closing total | 8 | 13.1 | 0.728 | 0.50 → 0.44 | Nothing beat chance |
| High-scoring (51+ and over the total) | 9 | 10.8 | 0.522 | 0.59 → 0.60 | Nothing beat chance |
| Over the closing total | 5 | 11.6 | 0.848 | 0.50 → 0.44 | Nothing beat chance |
| Upset (underdog wins outright) | 14 | 10.3 | 0.229 | 0.64 → 0.61 | Nothing beat chance |
| Underdog covers | 39 | 14.3 | 0.018 | 0.52 → 0.51 | 4 kept |
| FG-heavy (5+ made or 6+ tried) | 11 | 8.6 | 0.298 | 0.53 → 0.55 | Nothing beat chance |
| Run-heavy (300+ combined or a team 175+) | 152 | 12.5 | 0.001 | 0.44 → 0.67 | 3 kept |
| Pass-heavy (550+ combined or a team 325+) | 26 | 7.8 | 0.015 | 0.65 → 0.66 | 3 kept |
| Blowout (17+) | 14 | 8 | 0.159 | 0.60 → 0.64 | Nothing beat chance |
| One-score game (8 or less) | 23 | 11.2 | 0.070 | 0.58 → 0.56 | Nothing beat chance |
| Defensive or special-teams TD | 10 | 8.2 | 0.279 | 0.49 → 0.50 | Nothing beat chance |
| 4+ combined turnovers | 5 | 8.7 | 0.719 | 0.45 → 0.52 | Nothing beat chance |
| First half 13 or fewer points | 5 | 3.4 | 0.284 | 0.59 → 0.55 | Nothing beat chance |
| First half 28+ points | 12 | 11.6 | 0.378 | 0.59 → 0.60 | Nothing beat chance |
Popular ideas, tested the same way (thresholds on 2016–22 percentiles). None cleared the bar.
A second, wider search aimed at the rare ends of games: market structure (key numbers, moneyline vs spread, juice, opening vs closing line where history exists), referee crews, quarterback changes and backups, injured starters, rest, byes, travel, cold, wind, rain, altitude, late-season motivation, rematches, bounce-backs after extreme results, over/under and cover streaks, scoring luck, 4th-down aggressiveness and coach tenure — at extreme thresholds (top and bottom 2.5–10%) and in pairs.
The bar is lower than for stat trends, on purpose: at least 25 games in 2016–22, a lift of 8+ points over games on similar lines there, the same direction in 2023–24 and in the 2025 holdout, a football reason, and either the whole search beating shuffled results (permutation p ≤ 0.2) or a shrunk estimate still 4+ points above the base. That makes these small samples that may be noise. Only the best three per result show on game cards; each is demoted off the cards if its live record trails its base after 15 games.
77,100 rules tried across 12 results. Live records start Sep 27. Analysis and entertainment, not betting advice. 21+ only where legal.
Broad search: 6,425 rules, 26 passed the tail bar vs 39.3 on shuffled results (p 0.809, 1000 shuffles) · Theory list: 24 written-down hypotheses, 0 passed vs 0.05 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. Wind, rain, cold, backup quarterbacks, injured starters, under-leaning referee crews, under streaks, slow pace and conservative 4th-down coaches were all tested; rain/snow helps the plain under (see Under), but nothing moved the 37-or-fewer low-scoring result beyond chance.
Broad search: 6,425 rules, 51 passed the tail bar vs 64.5 on shuffled results (p 0.712, 1000 shuffles) · Theory list: 24 written-down hypotheses, 1 passed vs 0.07 (p 0.075) · family-wide shrinkage prior 4,866 games · holdout games in kept leans 15
When: rain or snow at an outdoor stadium (backtest: nflverse game-day weather; live: the kickoff forecast, 50%+ chance or rain/snow in the forecast).
Why it could be real: A wet ball means more fumbles and drops, shorter passes and more running, which runs the clock. Totals are set days ahead and move only partly with the forecast. Written down as a hypothesis before testing (theory family).
Shrunk estimate (100-game prior toward the base): 59% vs base 52% (+7 pts) · family-wide shrinkage: 52% · 2023–25 combined 28/43 vs 51% · theory-list permutation p 0.075
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 8 of 10 seasons beat the base.
Broad search: 6,425 rules, 7 passed the tail bar vs 14.7 on shuffled results (p 0.851, 1000 shuffles) · Theory list: 24 written-down hypotheses, 0 passed vs 0.01 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. The line alone already explains the 30-or-fewer games; every extra condition was at or below what shuffled results produce.
Broad search: 6,425 rules, 59 passed the tail bar vs 64.3 on shuffled results (p 0.545, 1000 shuffles) · Theory list: 20 written-down hypotheses, 0 passed vs 0.08 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. Dome + fast pace, over-leaning crews, depleted defenses, over streaks, aggressive coaches, heat, altitude, week 18 and scoring-luck regression were all tested: fewer rules passed on the real results than on shuffled ones.
Broad search: 6,425 rules, 63 passed the tail bar vs 69.5 on shuffled results (p 0.555, 1000 shuffles) · Theory list: 20 written-down hypotheses, 0 passed vs 0.10 (p 1.000) · family-wide shrinkage prior 4,866 games · holdout games in kept leans 0
Nothing, even at the tail level. Same inputs as high scoring; nothing beat the shuffles.
Broad search: 6,425 rules, 23 passed the tail bar vs 32.6 on shuffled results (p 0.706, 1000 shuffles) · Theory list: 20 written-down hypotheses, 0 passed vs 0.03 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. Nothing beat the shuffles; the 60+ games were no more common where any tested condition (or the variance model) said they should be.
Broad search: 6,425 rules, 78 passed the tail bar vs 52.7 on shuffled results (p 0.127, 1000 shuffles) · Theory list: 28 written-down hypotheses, 1 passed vs 0.15 (p 0.143) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 149
When: the favorite's defense allows 0.20+ fewer EPA per play than the underdog's, opponent-adjusted (biggest 5% of gaps); and it allows 5.6+ fewer points per game, opponent-adjusted (biggest 20%).
Why it could be real: Defensive numbers are much less stable week to week than offensive ones, yet the spread leans on them. A favorite whose whole edge is its defense is priced on something that tends to regress, so the underdog wins outright more often than the line implies.
Shrunk estimate (100-game prior toward the base): 35% vs base 26% (+9 pts) · family-wide shrinkage: 27% · 2023–25 combined 8/25 vs 27% · broad-search permutation p 0.127
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: the underdog is one of the most-followed franchises (Cowboys, Chiefs, Packers, Steelers, Patriots, 49ers, Eagles); and its quarterback started 15+ of the team's last 16 games.
Why it could be real: Popular teams are rarely underdogs; when they are, it is usually after a short bad stretch the market overreacts to. With their established starter they are closer to their true level than the price.
Shrunk estimate (100-game prior toward the base): 45% vs base 38% (+7 pts) · family-wide shrinkage: 38% · 2023–25 combined 24/48 vs 35% · broad-search permutation p 0.127
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 8 of 10 seasons beat the base.
When: the underdog's red-zone touchdown rate, averaged with what the favorite's defense allows, is 41.5% or lower (bottom 2.5%).
Why it could be real: Red-zone touchdown rate is one of the noisiest team stats and snaps back toward average. An underdog whose scoring is dragged down by settling for field goals is better than its points say, and the spread partly prices the points.
Shrunk estimate (100-game prior toward the base): 38% vs base 32% (+7 pts) · family-wide shrinkage: 32% · 2023–25 combined 17/33 vs 33% · broad-search permutation p 0.127
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: the two offenses average 33.4 or fewer points per game combined (bottom 2.5%).
Why it could be real: When neither side can score, a few plays (a turnover, a long field goal, a special-teams swing) decide the game and the talent gap the spread measures matters less.
Shrunk estimate (100-game prior toward the base): 43% vs base 37% (+6 pts) · family-wide shrinkage: 37% · 2023–25 combined 13/25 vs 40% · broad-search permutation p 0.127
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: both offenses are slow between snaps: 60.6+ seconds per play combined (slowest 10%).
Why it could be real: Fewer snaps mean fewer possessions and a noisier result: the better team has fewer chances for its edge to show. Written down as a hypothesis before testing (theory family).
Shrunk estimate (100-game prior toward the base): 39% vs base 34% (+5 pts) · family-wide shrinkage: 34% · 2023–25 combined 76/214 vs 32% · theory-list permutation p 0.143
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 9 of 10 seasons beat the base.
When: the underdog won its previous game this season by 25+ points (top 2.5%).
Why it could be real: A team that just routed someone is often better than the season-long numbers the line leans on, and the market moves only part of the way.
Shrunk estimate (100-game prior toward the base): 43% vs base 37% (+6 pts) · family-wide shrinkage: 38% · 2023–25 combined 16/28 vs 35% · broad-search permutation p 0.127
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 7 of 10 seasons beat the base.
Broad search: 6,425 rules, 34 passed the tail bar vs 33.5 on shuffled results (p 0.413, 1000 shuffles) · Theory list: 28 written-down hypotheses, 0 passed vs 0.08 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. Nothing beat the shuffles for underdog wins by 7+.
Broad search: 6,425 rules, 44 passed the tail bar vs 47.7 on shuffled results (p 0.510, 1000 shuffles) · Theory list: 15 written-down hypotheses, 0 passed vs 0.03 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. Kicker-friendly domes, stout red-zone defenses, conservative coaches, field-goal referee crews and long-kick teams were tested; nothing beat the shuffles for 5+ made / 6+ tried.
Broad search: 6,425 rules, 21 passed the tail bar vs 10.3 on shuffled results (p 0.090, 1000 shuffles) · Theory list: 15 written-down hypotheses, 0 passed vs 0.00 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 7
When: the two offenses average 10.5 or fewer yards per play combined (bottom 20%); and the matchup projects 55.9+ combined carries (top 10%).
Why it could be real: Run-first offenses that don't gain chunk yardage move the chains but stall between the 20s, and conservative play-callers kick there: drives end in field-goal range far more often than in the end zone.
Shrunk estimate (100-game prior toward the base): 15% vs base 9% (+5 pts) · family-wide shrinkage: 9% · 2023–25 combined 9/48 vs 10% · broad-search permutation p 0.090
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 8 of 10 seasons beat the base.
Broad search: 6,425 rules, 29 passed the tail bar vs 44.7 on shuffled results (p 0.820, 1000 shuffles) · Theory list: 18 written-down hypotheses, 0 passed vs 0.07 (p 1.000) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 0
Nothing, even at the tail level. Backup QBs, injured starters, eliminated or resting teams, week 18, travel, and collapsing underdogs were tested; fewer blowout rules passed than on shuffled results (p 0.82).
Broad search: 6,425 rules, 65 passed the tail bar vs 72.8 on shuffled results (p 0.583, 1000 shuffles) · Theory list: 15 written-down hypotheses, 1 passed vs 0.06 (p 0.059) · family-wide shrinkage prior 5,000 games · holdout games in kept leans 24
When: the two offenses average 37.2 or fewer points per game combined (bottom 10%).
Why it could be real: Low-scoring teams rarely build big leads, so more of these games are decided by one score than the spread alone suggests. Held in 2016–22 and 2023–24 but was flat in the 2025 holdout. Written down as a hypothesis before testing (theory family).
Shrunk estimate (100-game prior toward the base): 60% vs base 54% (+7 pts) · family-wide shrinkage: 54% · 2023–25 combined 63/103 vs 55% · theory-list permutation p 0.059
Bars: hit rate each season (green = beat its base, red = didn't); line: base on similar lines. 10 of 10 seasons beat the base.
Machine-learning tails. Gradient boosting and logistic regression on every pregame input, judged only on the games they rated most likely (top 5% by predicted rate minus the line base). A model can miss overall and still find a real extreme slice; these didn't: in 2016–22 (season-by-season cross-validation) the top 5% hit about the base rate for every result.
| Result | Top 5% · 2016–22 | 2023–24 | 2025 |
|---|---|---|---|
| Low scoring (37 or fewer, under the total) | 22/96 (23% vs 26%) | 1/7 (14% vs 34%) | 3/8 (38% vs 26%) |
| Under the total | 37/95 (39% vs 51%) | 13/31 (42% vs 41%) | 18/41 (44% vs 37%) |
| 30 or fewer total points | 10/96 (10% vs 15%) | 2/28 (7% vs 12%) | 0/15 (0% vs 9%) |
| High scoring (51+, over the total) | 29/96 (30% vs 30%) | 8/15 (53% vs 23%) | 2/7 (29% vs 32%) |
| Over the total | 42/95 (44% vs 46%) | 13/39 (33% vs 44%) | 1/2 (50% vs 46%) |
| 60+ total points | 14/96 (15% vs 18%) | 2/20 (10% vs 10%) | 0/1 (0% vs 19%) |
| Upset (underdog wins) | 38/96 (40% vs 34%) | 22/50 (44% vs 34%) | 10/38 (26% vs 27%) |
| Underdog wins by 7+ | 16/96 (17% vs 19%) | 5/33 (15% vs 11%) | 2/19 (11% vs 10%) |
| FG-heavy (5+ made or 6+ tried) | 23/96 (24% vs 25%) | 14/46 (30% vs 18%) | 0/14 (0% vs 16%) |
| 6+ field goals | 8/96 (8% vs 9%) | 3/34 (9% vs 9%) | 1/14 (7% vs 7%) |
| Blowout (17+ margin) | 19/96 (20% vs 24%) | 1/7 (14% vs 24%) | 1/2 (50% vs 18%) |
| One-score game | 46/96 (48% vs 44%) | 7/19 (37% vs 54%) | 3/8 (38% vs 38%) |
Score-distribution model. We turned each closing line into tail probabilities (30 or fewer points, 60+, underdog by 7+, blowout, one-score) from how far real games landed from their lines, then let a model widen or narrow that spread game by game from the team stats. Its predicted volatility did not track how wild games actually were (correlation about zero in every split), and games it called "fatter-tailed" were not.
Line movement. Opening lines exist only for 2016–21 here, so movement rules could be explored but never checked on later seasons; none are used.