| Metric | What it measures | How it's calculated | Why it's here |
|---|---|---|---|
| SP+ | Overall team quality: how many points better (or worse) than a perfectly average FBS team on a neutral field. Offense/defense sub-ratings are expected points scored/allowed vs. an average opponent — for defense, lower is better. | Bill Connelly's rating system: a regression blend of play-by-play efficiency, opponent adjustment, and tempo control. Preseason versions add returning production, recruiting, and transfer data. | The single best publicly available “how good, really” number. Opponent adjustment makes it the fair way to compare teams from different leagues — our schedule-caveat referee. |
| Returning production | The share of last season's play value that is back on this year's roster, overall and by facet (passing / receiving / rushing). | Sum of returning players' PPA divided by the team's total PPA from last season (CFBD). | Tells you how much of last year's stats still describe this year's team — the difference between a stale baseline and a live one. |
| The model | This project's reference score projection. | Expected points = SP+ offense + opponent SP+ defense − league-average offense rating, plus 2.5 home-field. Logged to predictions.csv every week. | Exists to be tracked, not trusted: backtested on ~880 games (2023–25) it loses to the betting market, so it's shown as reference only. A big model-vs-market gap is a signal the market knows something worth finding. |
| Win probability / expected wins | The chance Texas wins a given game, and the sum of those chances across the schedule (used in the preseason report). | Projected margin = SP+ gap ± 2.5 home-field, converted to a probability with a normal distribution whose spread comes from the model's backtested error (MAE 14.05 → σ ≈ 17.6). Expected wins = the probabilities summed. | Turns one schedule into one number that can be compared against the market's win total. Inherits the model's reference-only status — it's a shape of the season, never a pick. |
The core family. Every unit comparison in the reports reduces to some combination of these four.
| Metric | What it measures | How it's calculated | Why it's here |
|---|---|---|---|
| PPA / EPA (Predicted / Expected Points Added) | The scoring value a play (or player, or unit) added — the core currency of modern football analytics. | Every down-distance-field-position state has an expected point value; a play's PPA is the change in that value from before the snap to after. Averaged per play or summed per player/season. CFBD's PPA and the generic term EPA are the same idea. | Values every play by what it did to the scoreboard's expectation — 3 yards on 3rd-and-2 counts as the win it is; 3 yards on 3rd-and-8 doesn't. |
| Success rate | Consistency: the share of plays that keep the offense “on schedule.” | A play succeeds if it gains ≥50% of needed yards on 1st down, ≥70% on 2nd, or 100% on 3rd/4th. | The floor metric. Paired with explosiveness it diagnoses an offense's personality — high success + low explosiveness grinds; the reverse is boom/bust. |
| Explosiveness (IsoPPP) | The size of the successful plays — a unit's ceiling. | Average PPA on successful plays only (isolating magnitude from frequency). | Big plays decide games disproportionately. “Explosiveness allowed” is the big-play-problem stat for defenses. |
| Points per scoring opportunity | Finishing: how many points a drive produces once it reaches scoring range. | Points scored divided by drives that reached the opponent's 40-yard line. | Scoring efficiency without pace distortion — teams that move the ball but settle for field goals show up here. |
The stats built to answer “who earned that yard?” — separating linemen from linebackers from ball-carriers.
| Metric | What it measures | How it's calculated | Why it's here |
|---|---|---|---|
| Line yards | The offensive line's share of the running game. | Rushing yards credited to the line with distance weights: losses count 125%, yards 0–4 count 100%, 5–10 count 50%, beyond 10 count 0%. Per carry. | Splits credit between blocking and running — the stat that let us catch Texas's 2025 run-blocking problem hiding behind decent team rushing totals. |
| Second-level / open-field yards | Yards earned past the line: 5–10 out (second level, mostly on linebackers) and 10+ (open field, mostly the back's creation). | Per-carry yardage in each distance band (CFBD advanced rushing). | Completes the run-game split: line → linebackers → the back himself. |
| Stuff rate | How often runs die at or behind the line of scrimmage. | Stuffed carries ÷ total carries. Offense wants it low; defense wants it high. | A pure trench-win stat with no skill-player contamination. |
| Power success | Short-yardage muscle. | Conversion rate on 3rd/4th down with ≤2 yards to go (plus goal-to-go from the 2). | The must-have-a-yard situations that swing drives — an OL identity check. |
| Havoc rate | Disruption: how often a defense wrecks a play. | (Tackles for loss + forced fumbles + interceptions + passes broken up) ÷ opponent plays. Split into front-seven havoc and DB havoc. | The Simmons metric up front; in the secondary it separates ball-hawking units from passive ones (Texas State's 4.7% DB havoc was 14th percentile — the quantified fatal flaw). |
| Field position (avg start) | The hidden yardage of special teams, turnovers, and coverage units. | Average yard line where a team's offensive drives begin (measured as distance to the goal — lower is better). | Quietly worth points every game: an 85th-vs-14th percentile gap means Texas repeatedly starts drives a first down or two ahead. |
The early-warning family: numbers that predict their own reversal. Effect sizes here are backtested on 2023–25 (see BACKTEST.md in the repo).
| Metric | What it measures | How it's calculated | Why it's here |
|---|---|---|---|
| Turnover margin /game | Takeaways minus giveaways, per game. | (Opponent turnovers − own turnovers) ÷ games. | Backtested regression flag: in 2023–25, teams beyond +0.5/game lost ~9% of win percentage the next season; beyond −0.5 gained ~12%. Big margins are borrowed wins. |
| Fumble margin | The luckiest number in football. | Fumbles recovered − fumbles lost. | Recovering a live ball is close to a coin flip, so this regresses hardest toward zero — Texas's +5 was borrowed; Texas State's −7 was theft. |
| INT margin | Interceptions made minus thrown. | Team INTs − QB INTs thrown. | Part skill (QB decision-making, ball-hawking DBs), part bounce — treated as semi-durable, unlike fumble margin. |
| Base rates | What actually happened in games like this one. | From our committed 2023–25 snapshot: all favorites at or above a given spread, with straight-up, cover, and margin outcomes (tools/backtest.py --favorites). | Turns an abstract spread into calibrated history — e.g., 31.5+ favorites went 234-1 with a median margin of 41. |
Not metrics, but rules of the house — how numbers become bars, chips, and calls.
| Metric | What it measures | How it's calculated | Why it's here |
|---|---|---|---|
| Percentile | Where a value ranks nationally. | Share of all FBS teams (~136) that this value beats, direction-aware — a low defensive PPA scores a high percentile. 100% = best. | Raw values like “0.279 rushing PPA” mean nothing alone; percentiles make unlike metrics comparable at a glance. |
| Computed gap (Matchup Board bars) | The measured distance between two opposing units. | Offense unit's percentile minus the opposing defense unit's percentile on the pairing's headline metric; bar length = half the gap, filled toward the leader. Never hand-tuned. | Keeps the picture honest: data draws the bar, judgment only writes the chip — and must explain itself when they disagree. |
| Edge chip | The weekly report's call on a pairing after adjusting for schedule strength and roster change, on a fixed symmetric scale: Edge: Team, big · Edge: Team · Edge: Team, slight · Edge: Even. | Judgment, informed by SP+ (which is opponent-adjusted) and offseason movement — not computed. Tiers: big = this pairing alone can shape the scoring; plain = clear advantage; slight = margins, one roster change from flipping; Even = no call. Calibration guidance (not formula): slight ≈ adjusted gap under ~20 percentile points, big ≈ over ~45. Parentheses may scope a call, never intensify it; commentary lives in the note, never the chip. | Raw percentiles describe old rosters and unequal schedules; the chip is where that context lives — signed, visible, and in the same vocabulary every week. |
| Outlook chip (preseason) | The preseason report's season-level call for a unit, paired with its computed last-season percentile (stated as a number — no bar, since nothing is being compared). | Judgment on a fixed four-tier scale: Strength / Lean strength / Open question / Rebuilt. When the note says the players who earned the percentile departed, the number is history, not a forecast. | Same honesty contract as the edge chip: data states the number, judgment writes the chip, and disagreements get explained in the note. Bars are reserved for genuine two-unit comparisons (the weekly Matchup Board). |
| Garbage time | Excluded from all advanced stats. | Plays with the game effectively decided (large second-half leads) are dropped by CFBD's filter. | Keeps blowout mop-up drives from polluting efficiency numbers — especially important for big favorites. |
| Injury impact tiers | How an absence changes the Matchup Board. | The player's share of his team's PPA in the affected facet: ≥30% downgrades the unit's edge one tier; ≥60% two tiers (tools/injury_impact.py). | Converts “star questionable” headlines into a sized, repeatable adjustment. |
| Market terms | Spread / over-under / moneyline. | Spread: expected margin (a −7 spread = favored by 7 points). Over/under: expected total points. Moneyline: odds to win outright; −8000 implies ~98.8% (risk 8000 to win 100). | The market is the best public forecast — our backtest proved it beats our model — so its numbers anchor every verdict. |