Skip to main content

CFB — additional Python functions — Analytics: cfb_season–play_type

cfb_season_odds​

cfb_season_odds(seasons: 'int | list[int]', *, as_of_date: 'datetime.date | None' = None, n_sims: 'int' = 10000, playoff_seeds: 'int' = 12, seed: 'int' = 0, era: 'str' = 'modern', return_as_pandas: 'bool' = False) -> 'pl.DataFrame | pd.DataFrame'

Ratings-driven season Monte Carlo: conference / playoff / championship odds.

Thin wrapper over cfb_simulations.cfb_simulations -- it builds the ratings with cfb_ratings.cfb_ratings, converts the schedule to the engine format with cfb_standings.cfb_games_from_schedule (re-keyed on ESPN team_id so the ratings align), and feeds make_ratings_compute_results as the sampler. All season / standings / bracket machinery is reused; unplayed games are simulated, played games (before as_of_date) are kept. Only FBS programs (schedule division == "fbs") enter the simulated universe; non-FBS opponents stay in the game set -- their games still count toward FBS records -- but can never reach the standings, the playoff field, or the output.

Parameters

ParameterTypeDefaultDescription
seasonsint | list[int]A single season (an int, or a one-element list). Multiple seasons raise ValueError -- the simulation engine is single-season.
as_of_datedate | NoneNoneLeakage boundary applied to BOTH the ratings vintage and the game set. Ratings are fit only on plays from games with date < as_of_date (cfb_ratings.cfb_ratings), and schedule results from start_date on/after as_of_date are masked so those games are simulated instead of replayed; masked postseason rows are dropped (the matchup is itself an outcome) and regenerated from each sim's own standings. None uses the full season as-is.
n_simsint10000Number of simulated seasons.
playoff_seedsint12CFP field size.
seedint0RNG seed for reproducibility.
erastr'modern'Era key into cfb_prediction_constants.CFB_CONSTANTS.
return_as_pandasboolFalseIf True, return a pandas DataFrame; otherwise polars.

Returns

One row per team: season, team_id (Utf8), exp_wins, conf_title_prob, playoff_prob, first_round_bye_prob, cfp_champ_prob (Float64 probabilities in [0, 1]). Zero-row (typed) when no ratings/schedule are available.

col_nametypedescription
seasonintegerSeason simulated (null for a pooled multi-season call).
team_idcharacterTeam ESPN id (character join key).
exp_winsdoubleMean wins per simulated season.
conf_title_probdoubleShare of simulations in which the team won its conference.
playoff_probdoubleShare of simulations in which the team made the College Football Playoff field.
first_round_bye_probdoubleShare of simulations in which the team earned a CFP first-round bye.
cfp_champ_probdoubleShare of simulations in which the team won the College Football Playoff national championship.

Example

from sportsdataverse.cfb.cfb_season_odds import cfb_season_odds
odds = cfb_season_odds(2023, n_sims=2000)
odds.sort("cfp_champ_prob", descending=True).head()

cfb_transfer_impact​

cfb_transfer_impact(target_season: 'int | list[int]', *, division: 'str' = 'fbs', alpha: 'float' = 1.0, return_as_pandas: 'bool' = False) -> 'pl.DataFrame | pd.DataFrame'

Net transfer talent and its projected win-total impact per team-season.

pred_win_delta comes from an on-demand ridge of realized win deltas on net_transfer_talent fitted over strictly-prior seasons (the as-of boundary is enforced internally per target season).

Parameters

ParameterTypeDefaultDescription
target_seasonint | list[int]Season (or list) to score.
divisionstr'fbs'Division slug for the star-points constants.
alphafloat1.0Ridge L2 penalty.
return_as_pandasboolFalseIf True, return a pandas DataFrame; otherwise polars.

Returns

Per (season, team_id): net_transfer_talent (Float64), pred_win_delta (Float64). Zero-row (typed) when no data.

col_nametypedescription
seasonintegerSeason the net transfer talent describes.
team_idcharacterESPN team id as a string (rosters team_id, cast Int64 to Utf8).
net_transfer_talentdoubleIncoming minus outgoing transfer talent points for the season.
pred_win_deltadoubleRidge-projected win-total change from net transfer talent (as-of fit; weak observed validity - see the strict-xfail gate).

Example

from sportsdataverse.cfb import cfb_transfer_impact
imp = cfb_transfer_impact(2024)
imp.sort("net_transfer_talent", descending=True).head(10)

cfb_transfer_moves​

cfb_transfer_moves(seasons: 'int | list[int]', *, division: 'str' = 'fbs', return_as_pandas: 'bool' = False) -> 'pl.DataFrame | pd.DataFrame'

Transfer moves inferred from year-over-year roster diffs.

Parameters

ParameterTypeDefaultDescription
seasonsint | list[int]Destination season(s) to extract moves for (each compares S-1 -> S).
divisionstr'fbs'Division slug for the star-points constants.
return_as_pandasboolFalseIf True, return a pandas DataFrame; otherwise polars.

Returns

One row per move side: season (Int64, the destination season), team_id (Utf8 ESPN team id), player_id (Utf8 ESPN athlete id), direction ("in" | "out"), prior_team_id (Utf8 ESPN team id of the season S-1 team), talent_points (Float64, name-joined to the cfb_recruits release; the 0-star default when the player has no recruit rating). Zero-row (typed) when rosters are unavailable.

col_nametypedescription
seasonintegerDestination season of the move (compares rosters S-1 to S).
team_idcharacterESPN team id (as a string) of the side this row describes (destination for "in", origin for "out").
player_idcharacterESPN athlete id as a string.
directioncharacterMove side - "in" (arriving at team_id) or "out" (leaving team_id).
prior_team_idcharacterESPN team id (as a string) of the season S-1 team.
talent_pointsdoubleRecruit-star talent points (name-matched to the 247 recruit record; 0-star default when unrated).

Example

from sportsdataverse.cfb import cfb_transfer_moves
moves = cfb_transfer_moves(2024)
moves.filter(pl.col("direction") == "in").group_by("team_id").len()

create_drive_summary​

create_drive_summary(drives: list[dict] | dict, frame: polars.dataframe.frame.DataFrame, home_id: str | int, away_id: str | int, periods: set[int] | str | None = None) -> dict | None

Build the StatBroadcast-style drive summary, chart, and long-play lists.

A drive belongs to the quarter it STARTED in. On a windowed build the full drive sequence still provides context (running score, the previous drive for OBTAINED and points-off-turnovers), but only in-window drives are counted, charted, or listed. largest_lead and the time-leading / time-tied split window too: the score state is read from the whole regulation play sequence and then clipped to the window's clock intervals (one per contiguous run of quarters, so a gapped set never charges the quarter it skipped). Under "ot" the clock has no axis to integrate over and only largest_lead ships.

Parameters

ParameterTypeDefaultDescription
driveslist[dict]the ESPN drives grouping, in game order (previous plus the in-progress current drive, if any).
framepl.DataFramethe enriched plays frame from CFBPlayProcess.run_processing_pipeline (plays_frame).
home_idstr | intESPN home team id.
away_idstr | intESPN away team id.
periodsset[int] | str | NoneNoneoptional window -- a set of quarter numbers (e.g. {1, 2}) or the string "ot" (every period > 4). None = full game.

Returns

{"teams": {...}, "chart": [...], "scores": [...], "longPlays": {...}} keyed by team id, or None when the inputs are unusable (no drives, empty frame, or an empty window).

Example

summary = create_drive_summary(drives, game.plays_frame, "52", "61")

create_situational_stats​

create_situational_stats(frame: polars.dataframe.frame.DataFrame, home_id: str | int, away_id: str | int, window_expr: polars.expr.expr.Expr | None = None) -> dict | None

Build the situational team-stats block from a plays frame.

two_minute and middle_8 are omitted from a windowed build: both name a clock window of their own, so intersecting them with another window describes neither (middle-8 inside Q1 is empty). Every other section, pace and non_garbage included, is computed on the windowed slice and ships with it.

Parameters

ParameterTypeDefaultDescription
framepl.DataFramethe enriched plays frame from CFBPlayProcess.run_processing_pipeline (plays_frame).
home_idstr | intESPN home team id.
away_idstr | intESPN away team id.
window_exprExpr | NoneNoneoptional polars filter expression windowing the windowable sections to that slice (e.g. pl.col("period") == 3). None = full game.

Returns

{"teams": {<team_id>: {<section>: ...}}} or None when the frame is unusable or the window is empty.

Example

stats = create_situational_stats(game.plays_frame, "52", "61")

make_ratings_compute_results​

make_ratings_compute_results(ratings: 'pl.DataFrame', *, era: 'str' = 'modern') -> 'ComputeResultsFn'

Build a cfb_simulations compute_results closure from fixed ratings.

The returned closure implements the engine's results contract -- (teams, games, week_num, *, rng, **kwargs) -> {"teams", "games"} -- filling every unplayed week == week_num game's result with a sampled home margin round(Normal(exp_margin, margin_sd)), where exp_margin is cfb_game_predict.predict_margin on the two teams' adj_net (home-field applied unless neutral). Unlike the default elo sampler the ratings are fixed, so teams passes through unchanged (no elo update). Postseason games (game_type != "REG") re-break a sampled tie by win probability.

Parameters

ParameterTypeDefaultDescription
ratingsDataFrameA cfb_ratings.cfb_ratings-style frame with team_id and adj_net. Teams absent from it are treated as league-average (0.0).
erastr'modern'Era key into cfb_prediction_constants.CFB_CONSTANTS.

Returns

A compute_results callable suitable for cfb_simulations(..., compute_results=...).

Example

import numpy as np, polars as pl
from sportsdataverse.cfb.cfb_season_odds import make_ratings_compute_results
cr = make_ratings_compute_results(pl.DataFrame({"team_id": ["A", "B"], "adj_net": [0.3, -0.3]}))
teams = pl.DataFrame({"sim": [1, 1], "team": ["A", "B"], "conference": ["X", "X"]})
games = pl.DataFrame({"sim": [1], "week": [1], "home_team": ["A"], "away_team": ["B"],
"neutral": [0], "result": [None]})
cr(teams, games, 1, rng=np.random.default_rng(0))["games"]

play_type_family_expr​

play_type_family_expr(source: 'str' = 'play_type_canonical') -> 'pl.Expr'

Build the polars expression mapping a canonical type to its phase family.

Parameters

ParameterTypeDefaultDescription
sourcestr'play_type_canonical'Name of the canonical play-type column.

Returns

A pl.Expr aliased play_type_family; unmapped values yield null.

Example

import polars as pl
from sportsdataverse.cfb import add_play_type_canonical

pbp = pl.DataFrame({"type.text": ["Rush", "Timeout"]})
add_play_type_canonical(pbp).filter(
pl.col("play_type_family") != "administrative"
)