Skip to main content
Version: 0.1.5

MLB — additional Python functions — Models and calculators

as_of_split​

as_of_split(events: "'pl.DataFrame'", cutoff_date: 'Any', *, date_col: 'str' = 'game_date') -> "'pl.DataFrame'"

Leakage boundary: rows strictly before cutoff_date only.

The predictive path of the stolen-base (and, where predictive, baserunning) model must derive runner/catcher features only from data known before the event being scored -- this helper is the one place that boundary is enforced, so every predictive caller shares it.

Parameters

ParameterTypeDefaultDescription
eventsDataFrameAny frame carrying a date column.
cutoff_dateAnyExclusive upper bound (rows with date_col < cutoff_date are kept).
date_colstr'game_date'Name of the date column. Defaults to "game_date".

Returns

the filtered frame (unchanged if empty or missing date_col).

col_nametypedescription
pitch_typecharacterAbbreviation of the pitch type thrown (e.g. FF, SL, CH).
game_datecharacterGame date (YYYY-MM-DD).
release_speeddoublePitch velocity out of the hand (mph).
release_pos_xdoubleHorizontal release position of the ball, catcher's perspective (feet).
release_pos_zdoubleVertical release position of the ball, catcher's perspective (feet).
player_namecharacterPlayer name.
batterintegerMLBAM player id of the batter.
pitcherintegerWhether the position is a pitcher.
eventscharacterNested list of non-game events.
descriptioncharacterLong-form description text.
spin_dirdoubleDeprecated spin direction field, no longer populated.
spin_rate_deprecateddoubleDeprecated legacy spin-rate field, no longer populated.
break_angle_deprecateddoubleDeprecated legacy break-angle field, no longer populated.
break_length_deprecateddoubleDeprecated legacy break-length field, no longer populated.
zonedoubleStrike-zone region the pitch crossed (1-14 Gameday zone).
descharacterFull text description of the play.
game_typecharacterGame type code (R, P, etc.).
standcharacterSide of the plate the batter is standing (L or R).
p_throwscharacterHand the pitcher throws with (L or R).
home_teamcharacterHome team name.
away_teamcharacterAway team name.
typecharacterRecord type / category.
hit_locationdoubleFielder position number that fielded the ball.
bb_typecharacterBatted-ball type (ground_ball, line_drive, fly_ball, popup).
ballsintegerBall count before the pitch.
strikesintegerStrike count before the pitch.
game_yearintegerSeason year of the game.
pfx_xdoubleHorizontal pitch movement from the catcher's perspective (feet).
pfx_zdoubleVertical pitch movement from the catcher's perspective (feet).
plate_xdoubleHorizontal position of the pitch crossing the plate (feet from center).
plate_zdoubleVertical position of the pitch crossing the plate (feet above ground).
on_3bintegerMLBAM ID of the runner on third base, if any.
on_2bintegerMLBAM ID of the runner on second base, if any.
on_1bintegerMLBAM ID of the runner on first base, if any.
outs_when_upintegerNumber of outs when the batter came to the plate.
inningintegerInning number.
inning_topbotcharacterHalf of the inning (Top or Bot).
hc_xdoubleHit coordinate X on the field diagram.
hc_ydoubleHit coordinate Y on the field diagram.
tfs_deprecateddoubleDeprecated time-from-start field, no longer populated.
tfs_zulu_deprecateddoubleDeprecated Zulu time-from-start field, no longer populated.
umpiredoubleDeprecated umpire field, no longer populated.
sv_iddoubleDeprecated Sportvision/Statcast pitch identifier, no longer populated.
vx0doubleVelocity of the pitch in the x-direction at y=50 ft (ft/s).
vy0doubleVelocity of the pitch in the y-direction at y=50 ft (ft/s).
vz0doubleVelocity of the pitch in the z-direction at y=50 ft (ft/s).
axdoubleAcceleration of the pitch in the x-direction at y=50 ft (ft/s^2).
aydoubleAcceleration of the pitch in the y-direction at y=50 ft (ft/s^2).
azdoubleAcceleration of the pitch in the z-direction at y=50 ft (ft/s^2).
sz_topdoubleTop of the batter's strike zone for the pitch (feet).
sz_botdoubleBottom of the batter's strike zone for the pitch (feet).
hit_distance_scdoubleStatcast-measured projected distance of the batted ball (feet).
launch_speeddoubleExit velocity of the batted ball (mph).
launch_angledoubleVertical launch angle of the batted ball (degrees).
effective_speeddoublePerceived velocity adjusted for release extension (mph).
release_spin_ratedoubleSpin rate of the pitch at release (rpm).
release_extensiondoubleDistance toward the plate at release (feet).
game_pkintegerUnique game identifier.
fielder_2integerMLBAM ID of the catcher.
fielder_3integerMLBAM ID of the first baseman.
fielder_4integerMLBAM ID of the second baseman.
fielder_5integerMLBAM ID of the third baseman.
fielder_6integerMLBAM ID of the shortstop.
fielder_7integerMLBAM ID of the left fielder.
fielder_8integerMLBAM ID of the center fielder.
fielder_9integerMLBAM ID of the right fielder.
release_pos_ydoubleRelease position of the ball toward the plate (feet).
estimated_ba_using_speedangledoubleExpected batting average based on exit velocity and launch angle.
estimated_woba_using_speedangledoubleExpected wOBA based on exit velocity and launch angle.
woba_valuedoublewOBA value assigned to the event.
woba_denomdoublewOBA denominator (plate-appearance weight) for the event.
babip_valuedoubleBABIP value assigned to the event (0 or 1).
iso_valuedoubleIsolated power value assigned to the event.
launch_speed_angledoubleBatted-ball classification code (1-6) from exit velocity and angle.
at_bat_numberintegerSequential plate-appearance number within the game.
pitch_numberintegerPitch number within the plate appearance.
pitch_namecharacterFull name of the pitch type (e.g. 4-Seam Fastball, Slider).
home_scoreintegerHome team run total after the play.
away_scoreintegerAway team run total after the play.
bat_scoreintegerBatting team score before the pitch.
fld_scoreintegerFielding team score before the pitch.
post_away_scoreintegerAway team score after the pitch.
post_home_scoreintegerHome team score after the pitch.
post_bat_scoreintegerBatting team score after the pitch.
post_fld_scoreintegerFielding team score after the pitch.
if_fielding_alignmentcharacterInfield defensive alignment (Standard, Strategic, Infield shift).
of_fielding_alignmentcharacterOutfield defensive alignment (Standard, Strategic, 4th outfielder).
spin_axisdoubleSpin axis of the pitch as a clock-face angle (degrees).
delta_home_win_expdoubleChange in home team win expectancy on the play.
delta_run_expdoubleChange in run expectancy on the play.
bat_speeddoubleBat speed at the point of contact (mph).
swing_lengthdoubleLength of the swing path to contact (feet).
miss_distancedouble
estimated_slg_using_speedangledoubleExpected slugging based on exit velocity and launch angle.
delta_pitcher_run_expdoubleChange in run expectancy credited to the pitcher.
hyper_speeddoubleAdjusted (90th-percentile) exit velocity (mph).
home_score_diffintegerHome team score minus away team score before the pitch.
bat_score_diffintegerBatting team score minus fielding team score before the pitch.
home_win_expdoubleHome team win expectancy before the play.
bat_win_expdoubleBatting team win expectancy before the play.
age_pit_legacyintegerPitcher age using the legacy calculation.
age_bat_legacyintegerBatter age using the legacy calculation.
age_pitintegerPitcher age for the season.
age_batintegerBatter age for the season.
n_thruorder_pitcherintegerTimes through the order the pitcher is facing the lineup.
n_priorpa_thisgame_player_at_batintegerNumber of prior plate appearances by the batter in the game.
pitcher_days_since_prev_gamedoubleDays since the pitcher's previous game appearance.
batter_days_since_prev_gameintegerDays since the batter's previous game appearance.
pitcher_days_until_next_gamedoubleDays until the pitcher's next game appearance.
batter_days_until_next_gamedoubleDays until the batter's next game appearance.
api_break_z_with_gravitydoubleVertical pitch break including gravity (inches).
api_break_x_armdoubleHorizontal pitch break to the pitcher's arm side (inches).
api_break_x_batter_indoubleHorizontal pitch break toward/away from the batter (inches).
arm_angledoublePitcher's arm angle at release (degrees).
attack_angledoubleAngle of the bat's path at contact (degrees).
attack_directiondoubleHorizontal direction of the swing at contact (degrees).
swing_path_tiltdoubleVertical tilt of the swing path (degrees).
intercept_ball_minus_batter_pos_x_inchesdoubleHorizontal offset of ball-bat intercept from batter position (inches).
intercept_ball_minus_batter_pos_y_inchesdoubleDepth offset of ball-bat intercept from batter position (inches).

Example

from sportsdataverse.mlb.mlb_run_values import as_of_split
history = as_of_split(events, cutoff_date=dt.date(2024, 6, 15))

build_we_table​

build_we_table(states: 'pl.DataFrame', results: 'pl.DataFrame', *, laplace: 'float' = 1.0) -> 'pl.DataFrame'

Empirical, Laplace-smoothed home win-expectancy table.

Parameters

ParameterTypeDefaultDescription
statesDataFrameOutput of pbp_base_out_states.
resultsDataFrameGame-level results with game_id (same dtype as states), home_score, away_score.
laplacefloat1.0Additive smoothing constant (default 1.0).

Returns

one row per observed state bucket. | Column | Type | Description | |---|---|---| | inning_capped | Int64 | Inning, capped at 9 | | half | Utf8 | "top" or "bottom" | | base_state | Utf8 | 3-char base occupancy | | outs_start | Int64 | Outs before the play (0-2) | | score_diff_bucket | Int64 | home - away score, clipped to [-6, 6] | | home_win_exp | Float64 | Laplace-smoothed P(home wins | state) | | n | Int64 | Plate appearances observed in this bucket |

col_nametypedescription
inning_cappedintegerInning number, capped at 9 (extra innings pooled with the 9th).
halfcharacterHalf-inning ("top" or "bottom").
base_statecharacter3-char base occupancy code.
outs_startintegerOuts before the play (0-2).
score_diff_bucketintegerhome minus away score, clipped to [-6, 6].
home_win_expdoubleLaplace-smoothed empirical P(home team wins | state bucket).
nintegerPlate appearances observed in this state bucket.

Example

from sportsdataverse.mlb.mlb_run_expectancy import pbp_base_out_states
from sportsdataverse.mlb.mlb_win_expectancy import build_we_table
states = pbp_base_out_states(pbp)
table = build_we_table(states, results)

count_strike_run_value​

count_strike_run_value(pitches: "'pl.DataFrame'") -> "'pl.DataFrame'"

Ball-to-strike run-expectancy delta per count, from delta_run_exp.

strike_run_value is positive = runs saved by the defense per stolen strike, since a called strike carries negative delta_run_exp for the batting team relative to a ball in the same count: strike_run_value = -(E[delta_run_exp | called_strike, count] - E[delta_run_exp | ball, count]).

Parameters

ParameterTypeDefaultDescription
pitchesDataFramePitch-level frame with balls, strikes, description, and delta_run_exp columns (a sportsdataverse.mlb.mlb_statcast_extra.mlb_statcast_search frame). Rows other than called_strike/ball are ignored.

Returns

one row per observed count. | Column | Type | Description | |---|---|---| | balls | Int64 | Ball count (0-3) entering the pitch | | strikes | Int64 | Strike count (0-2) entering the pitch | | strike_run_value | Float64 | Runs saved by the defense per called strike vs. a ball in this count |

col_nametypedescription
ballsintegerBall count before the pitch.
strikesintegerStrike count before the pitch.
strike_run_valuedouble

Example

from sportsdataverse.mlb.mlb_run_values import count_strike_run_value
rv = count_strike_run_value(pitches)

event_run_value​

event_run_value(pitches: "'pl.DataFrame'", events: "'List[str]'") -> 'float'

Empirical run value of an event set, from mean delta_run_exp.

Parameters

ParameterTypeDefaultDescription
pitchesDataFramePitch-level frame with an events column and delta_run_exp.
eventsList[str]Statcast events values to average over (e.g. ["stolen_base_2b"]).

Returns

Mean delta_run_exp over rows whose events is in events. 0.0 if the frame is empty, lacks delta_run_exp, or no rows match.

Example

from sportsdataverse.mlb.mlb_run_values import event_run_value
rv_sb = event_run_value(pitches, ["stolen_base_2b", "stolen_base_3b"])

mae​

mae(a: 'np.ndarray', b: 'np.ndarray') -> 'float'

Mean absolute error between two arrays.

Parameters

ParameterTypeDefaultDescription
andarrayFirst array of values.
bndarraySecond array of values (same length as a).

Returns

The mean absolute error.

Example

import numpy as np
from sportsdataverse._common.metrics import mae
mae(np.array([1.0, 2.0]), np.array([1.5, 2.5]))

mlb_batter_projection​

mlb_batter_projection(target_season: 'int', *, history: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Next-season xwOBA projection (Marcel + delta-method aging) for every batter.

If history is None, builds player-season xwOBA history via sportsdataverse.mlb.mlb_expected_stats.mlb_expected_stats across the three seasons before target_season (ages must already be present on a supplied history frame -- this convenience path is intended for callers who already maintain an age-joined roster history).

Parameters

ParameterTypeDefaultDescription
target_seasonintThe season being projected.
historyOptional[DataFrame]NonePre-built player-season history (batter, season, age, xwoba, pa). If None, uses mlb_expected_stats over target_season - 3 .. target_season - 1.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per batter: age, proj_xwoba, proj_pa. Empty history returns a zero-row frame with the documented schema.

col_nametypedescription
batterintegerMLBAM batter id.
ageintegerBatter's projected age in target_season (last known age plus one).
proj_xwobadoubleMarcel-style weighted, PA-regressed, aging-curve-adjusted xwOBA projection for target_season.
proj_padoubleSum of plate appearances across the weighted lookback seasons used to build the projection.

Example

from sportsdataverse.mlb.mlb_batter_projection import mlb_batter_projection

proj = mlb_batter_projection(2024, history=player_season_history)
print(proj.shape)

# Pipeline next step (one line)

proj.sort("proj_xwoba", descending=True).head()

mlb_command_plus​

mlb_command_plus(pitches: 'pl.DataFrame', *, level: 'str' = 'pitch', return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Score pitches with the bundled Command+/Location+ (②) run-value model.

Parameters

ParameterTypeDefaultDescription
pitchesDataFrameOutput of sportsdataverse.mlb.mlb_pitch_features.pitch_features (needs plate_x_abs, plate_z_norm, in_zone, dist_from_heart, balls, strikes, stand, p_throws, pitch_type).
levelstr'pitch'"pitch" (default) for per-pitch output, or "pitcher" for a per-pitcher mean.
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

level="pitch": pitcher, pitch_type, location_rv_hat, command_plus. level="pitcher": pitcher, location_rv_hat, command_plus (per-pitcher mean). Empty input returns a zero-row frame with the documented schema.

col_nametypedescription
pitcherintegerMLB Advanced Media (MLBAM) id for the pitcher.
pitch_typecharacterStatcast pitch-type abbreviation.
location_rv_hatdoublePredicted per-pitch run value from the bundled Command+/Location+ xgboost model (location + count/handedness/pitch-type features only).
command_plusdoublePlus-scale Command+/Location+ score, 100 = league average, higher = better.

Example

from sportsdataverse.mlb.mlb_pitch_features import pitch_features
from sportsdataverse.mlb.mlb_command_plus import mlb_command_plus
feats = pitch_features(raw_pitches)
out = mlb_command_plus(feats)
print(out.select("pitcher", "command_plus").head())

# Pipeline next step

out.group_by("pitcher").agg(pl.col("command_plus").mean()).sort("command_plus", descending=True)

mlb_expected_home_runs​

mlb_expected_home_runs(start_dt: 'str', end_dt: 'str', *, puller: 'Optional[Callable[..., pl.DataFrame]]' = None, park_factors: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Per player-season park-neutral xHR, park-adjusted xHR, and HR-above-expected.

Pulls batted balls via puller(start_dt, end_dt, player_type="batter"), builds the EV x LA x spray HR-probability grid from the pull's own batted balls (season-agnostic algorithm, per-pull empirical constants), predicts each ball's HR probability with the EV x LA-marginal fallback, park-adjusts via hr_factor (index 100 = neutral, joined on the Statcast home_team abbreviation -> MLBAM team id), then aggregates per batter.

Parameters

ParameterTypeDefaultDescription
start_dtstrPull start date, YYYY-MM-DD.
end_dtstrPull end date, YYYY-MM-DD.
pullerOptional[Callable[..., DataFrame]]NoneInjectable Statcast search callable -- defaults to sportsdataverse.mlb.mlb_statcast_extra.mlb_statcast_search.
park_factorsOptional[DataFrame]NonePre-fetched park-factors frame (team_id, hr_factor); if None, fetched via sportsdataverse.mlb.mlb_statcast.mlb_statcast_leaderboard_park_factors.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (batter, season): hr, xhr_neutral, xhr_park_adj, hr_above_expected (hr - xhr_neutral). Empty pull returns a zero-row frame with the documented schema.

col_nametypedescription
batterintegerMLBAM batter id (join key into Savant's home-runs leaderboard as player_id).
seasonintegerFour-digit season year derived from game_year/game_date.
hrintegerRealized home runs in the pulled window.
xhr_neutraldoublePark-neutral expected home runs from the EV x LA x spray probability grid.
xhr_park_adjdoublePark-adjusted expected home runs (xhr_neutral cell probabilities scaled by each batted ball's home-park HR factor / 100).
hr_above_expecteddoublehr minus xhr_neutral -- positive means the batter over-performed the park-neutral HR model.

Example

from sportsdataverse.mlb.mlb_expected_home_runs import mlb_expected_home_runs

df = mlb_expected_home_runs("2024-06-01", "2024-06-21")
print(df.shape)

# Pipeline next step (one line)

df.sort("hr_above_expected", descending=True).head()

mlb_expected_stats​

mlb_expected_stats(start_dt: 'str', end_dt: 'str', *, puller: 'Optional[Callable[..., pl.DataFrame]]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Per player-season xwOBA/xBA/xSLG from an on-the-fly EV x LA empirical grid.

Pulls pitches via puller(start_dt, end_dt, player_type="batter"), builds the outcome grid from the pull's own batted balls (season-agnostic algorithm, per-pull empirical constants -- see CLAUDE.md), predicts contact woba/ba/slg per batted ball with the launch-angle- marginal fallback, then aggregates:

  • xwoba = (sum(predicted_woba over balls in play) + sum(woba_value over non-batted-ball PA-ENDING outcomes)) / derived_woba_denom -- the denominator is DERIVED from events (PA enders minus intentional walks / sac bunts / catcher interference), never trusted from a cache vintage's woba_denom column. The numerator excludes those same zero-denominator events, and a PA-ending walk/HBP whose woba_value is null in a given vintage is filled with the fixed weights .69 / .72.
  • xba = (sum(predicted_ba over TRACKED at-bat balls in play) + sum(realized hits over UNTRACKED ones)) / ab -- a ball in play with no launch data cannot be predicted from the grid, so it takes its realized outcome exactly as xwoba does, rather than counting in ab with a zero numerator (which deflated league-mean xBA by the untracked share).
  • xslg -- same construction on total bases.
  • woba / ba -- the OBSERVED counterparts on the same denominators, so xwoba - woba is a luck-vs-skill delta needing no second source.

pa counts PLATE-APPEARANCE-ENDING rows only (events non-null), never raw pitches -- a Statcast search pull carries every pitch, and counting them (the pre-fix behavior) inflated pa/ab by ~4x and corrupted xba/xslg scales.

Parameters

ParameterTypeDefaultDescription
start_dtstrPull start date, YYYY-MM-DD.
end_dtstrPull end date, YYYY-MM-DD.
pullerOptional[Callable[..., DataFrame]]NoneInjectable Statcast search callable -- defaults to sportsdataverse.mlb.mlb_statcast_extra.mlb_statcast_search.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (batter, season): pa, ab, xwoba, xba, xslg, plus the observed woba / ba. Empty pull returns a zero-row frame with the documented schema.

col_nametypedescription
batterintegerMLBAM batter id (join key into Savant's expected-stats leaderboard as player_id).
seasonintegerFour-digit season year derived from game_year/game_date.
paintegerPlate appearances in the pulled window.
abintegerAt-bats (PA minus walks, HBP, and sacrifices) in the pulled window.
xwobadoubleExpected wOBA from the EV x LA empirical grid (contact) plus realized non-contact outcome value, divided by wOBA denominator.
xbadoubleExpected batting average from the EV x LA empirical grid's hit-indicator cell means on tracked balls in play, plus the realized hit indicator on balls in play carrying no launch data, over at-bats.
xslgdoubleExpected slugging percentage from the EV x LA empirical grid's total-bases cell means on tracked balls in play, plus realized total bases on balls in play carrying no launch data, over at-bats.
wobadoubleObserved wOBA over the same denominator as xwoba, so xwoba - woba is the batter's contact-quality luck gap without needing a second source.
badoubleObserved batting average over the same at-bat denominator as xba, so xba - ba is the batted-ball luck gap without needing a second source.

Example

from sportsdataverse.mlb.mlb_expected_stats import mlb_expected_stats

df = mlb_expected_stats("2024-06-01", "2024-06-21")
print(df.shape)

# Pipeline next step (one line)

df.sort("xwoba", descending=True).head()

mlb_pitch_classify​

mlb_pitch_classify(pitches: 'pl.DataFrame', *, max_components: 'int' = 6, seed: 'int' = 0, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Per-pitcher Gaussian-mixture pitch reclassification.

Parameters

ParameterTypeDefaultDescription
pitchesDataFrameOutput of sportsdataverse.mlb.mlb_pitch_features.pitch_features (needs velo_z, spin_z, pfx_x_z, pfx_z_z).
max_componentsint6Cap on GMM components considered per pitcher (BIC picks the best 1..min(max_components, n_pitch_types)).
seedint0Random seed for reproducible cluster labels.
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

pitcher, pitch_type, pitch_type_reclass, reclass_confidence (max posterior cluster responsibility). Pitchers with fewer than MIN_PITCHES_FOR_CLUSTERING pitches pass through the Savant label unchanged with reclass_confidence = 1.0. Empty input returns a zero-row frame with this schema.

col_nametypedescription
pitcherintegerMLB Advanced Media (MLBAM) id for the pitcher.
pitch_typecharacterSavant-reported pitch-type abbreviation.
pitch_type_reclasscharacterReclassified pitch-type label from the per-pitcher GMM clustering (may differ from pitch_type).
reclass_confidencedoubleMax posterior cluster responsibility (1.0 for low-volume pitchers passed through unchanged).

Example

from sportsdataverse.mlb.mlb_pitch_features import pitch_features
from sportsdataverse.mlb.mlb_pitch_classify import mlb_pitch_classify
feats = pitch_features(raw_pitches)
out = mlb_pitch_classify(feats, seed=0)
print(out.filter(out["pitch_type"] != out["pitch_type_reclass"]).head())

mlb_prop_strikeouts​

mlb_prop_strikeouts(team_k9: 'float', opp_k_rate: 'float', lg_k_rate: 'float', *, innings: 'float' = 9.0) -> 'float'

Expected pitcher/team strikeouts via a K/9-and-opponent-K-rate blend.

team_k9 / 9 * innings * (opp_k_rate / lg_k_rate).

Parameters

ParameterTypeDefaultDescription
team_k9floatTeam/pitcher strikeouts per 9 innings pitched.
opp_k_ratefloatOpponent's own strikeout rate (K per PA).
lg_k_ratefloatLeague-average strikeout rate.
inningsfloat9.0Innings pitched in this outing (default 9.0).

Returns

expected strikeouts.

Example

from sportsdataverse.mlb.mlb_prop_projection import mlb_prop_strikeouts
mlb_prop_strikeouts(9.0, 0.22, 0.22)

mlb_prop_team_runs​

mlb_prop_team_runs(home_off: 'float', away_def: 'float', lg_rpg: 'float', *, park_factor: 'float' = 1.0) -> 'float'

Expected team runs via a log5-style rate blend.

lg_rpg * (home_off / lg_rpg) * (away_def / lg_rpg) * park_factor.

Parameters

ParameterTypeDefaultDescription
home_offfloatTeam's own runs-scored-per-game rate.
away_deffloatOpponent's runs-allowed-per-game rate.
lg_rpgfloatLeague-average runs-per-game rate.
park_factorfloat1.0Park run-scoring multiplier (default neutral 1.0; a real park-factor table is a documented follow-on).

Returns

expected runs for the team in this matchup.

Example

from sportsdataverse.mlb.mlb_prop_projection import mlb_prop_team_runs
mlb_prop_team_runs(5.5, 5.0, 4.5)

mlb_props​

mlb_props(matchups: 'pl.DataFrame', ratings: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Expected team runs + strikeouts for a slate of matchups.

Parameters

ParameterTypeDefaultDescription
matchupsDataFrameOne row per game: game_id, home_team_id, away_team_id.
ratingsDataFramePer-team as-of-date rate table: team_id, off_rpg (runs scored/game), def_rpg (runs allowed/game), and optionally k9 + k_rate (see the module docstring -- strikeout columns are null without them). team_id must share a dtype with matchups' team-id columns.
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

one row per matchup. | Column | Type | Description | |---|---|---| | game_id | Utf8 | Game identifier | | home_team_id | Utf8 | Home team identifier | | away_team_id | Utf8 | Away team identifier | | exp_runs_home | Float64 | Expected home-team runs | | exp_runs_away | Float64 | Expected away-team runs | | exp_strikeouts_home | Float64 | Expected home-pitcher strikeouts (null if ratings lacks k9/k_rate) | | exp_strikeouts_away | Float64 | Expected away-pitcher strikeouts (null if ratings lacks k9/k_rate) |

col_nametypedescription
game_idcharacterGame identifier.
home_team_idcharacterHome team identifier.
away_team_idcharacterAway team identifier.
exp_runs_homedoubleExpected home-team runs (log5-style rate blend).
exp_runs_awaydoubleExpected away-team runs (log5-style rate blend).
exp_strikeouts_homedoubleExpected home-pitcher strikeouts (null when the ratings input lacks k9/k_rate).
exp_strikeouts_awaydoubleExpected away-pitcher strikeouts (null when the ratings input lacks k9/k_rate).

Example

from sportsdataverse.mlb.mlb_prop_projection import mlb_props
props = mlb_props(matchups, ratings)

mlb_pythagenpat​

mlb_pythagenpat(runs_scored: 'float', runs_allowed: 'float', games: 'int', *, exponent: 'float' = 0.287) -> 'float'

Pythagenpat expected win percentage (Smyth-Patriot, run-environment adaptive exponent).

x = ((runs_scored + runs_allowed) / games) ** exponent; win_pct = runs_scored**x / (runs_scored**x + runs_allowed**x).

Parameters

ParameterTypeDefaultDescription
runs_scoredfloatTotal runs scored.
runs_allowedfloatTotal runs allowed.
gamesintGames played.
exponentfloat0.287Run-environment exponent (default the published 0.287).

Returns

expected win percentage in [0, 1]. Returns 0.5 when games == 0 or runs_scored + runs_allowed == 0 (guard against a zero-division/degenerate input).

Example

from sportsdataverse.mlb.mlb_team_projection import mlb_pythagenpat
mlb_pythagenpat(800, 600, 162)

mlb_pythagenpat_table​

mlb_pythagenpat_table(results: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Per-(season, team) pythagenpat table from game-level results.

This is a same-window estimator, not a forward-looking prediction: pythagenpat smooths a team's already-known run differential into an implied "true-talent" win rate over that same window (the classic Bill James validation is exactly "does the formula's win% track the actual win% over the same season"). To use it predictively for a future game, pre-filter results to games strictly before that date with sportsdataverse.mlb.mlb_game_state_constants.as_of_split first -- this function does not do that filtering itself.

Parameters

ParameterTypeDefaultDescription
resultsDataFrameGame-level results (season, home_team_id, away_team_id, home_score, away_score).
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

one row per (season, team). | Column | Type | Description | |---|---|---| | season | Int64 | Season | | team_id | Utf8 | Team identifier | | runs_scored | Int64 | Total runs scored | | runs_allowed | Int64 | Total runs allowed | | games | Int64 | Games played | | win_pct | Float64 | Realized win percentage | | pythag_win_pct | Float64 | Pythagenpat expected win percentage |

col_nametypedescription
seasonintegerMLB season (4-digit start year).
team_idcharacterTeam identifier (statsapi team id, stringified).
runs_scoredintegerTotal runs scored across the covered games.
runs_allowedintegerTotal runs allowed across the covered games.
gamesintegerCount of games played in the season.
win_pctdoubleRealized win percentage.
pythag_win_pctdoublePythagenpat expected win percentage (exponent 0.287).

Example

from sportsdataverse.mlb.mlb_team_projection import mlb_pythagenpat_table
table = mlb_pythagenpat_table(results)

mlb_run_expectancy_matrix​

mlb_run_expectancy_matrix(seasons: 'Union[int, List[int], None]' = None, *, pbp: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Empirical RE24 run-expectancy matrix by base-out state.

re[base_state, outs] = mean(runs_rest_of_inning) over all plate appearances starting in that state, excluding the bottom of the 9th inning and beyond (the standard RE24 exclusion -- those half-innings are only played while the home team trails or is tied, a score-differential selection bias that would otherwise distort the matrix). Computed on demand from statsapi play-by-play; no bundled artifact.

Parameters

ParameterTypeDefaultDescription
seasonsUnion[int, List[int], None]NoneOne season (int) or a list of seasons to collect via sportsdataverse.mlb.mlb_api_extra.mlb_schedule. Ignored when pbp is supplied.
pbpOptional[DataFrame]NonePre-collected parsed play-by-play frame (skips the network collector -- primarily for tests / offline reuse).
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

up to 24 rows (base_state x outs). | Column | Type | Description | |---|---|---| | base_state | Utf8 | 3-char base occupancy (e.g. "1_3") | | outs | Int64 | Outs at the start of the state (0-2) | | re | Float64 | Mean runs scored through the end of the half-inning | | n | Int64 | Number of plate appearances observed in this state |

col_nametypedescription
base_statecharacter3-char base occupancy code ("_" = empty, "1"/"2"/"3" = occupied), e.g. "1_3" for runners on first and third.
outsintegerOuts at the start of the base-out state (0-2).
redoubleEmpirical mean runs scored from this state through the end of the half-inning (RE24).
nintegerNumber of plate appearances observed starting in this base-out state.

Example

from sportsdataverse.mlb.mlb_run_expectancy import mlb_run_expectancy_matrix
matrix = mlb_run_expectancy_matrix(pbp=pbp)

# Pipeline next step (one line)

matrix.filter(pl.col("base_state") == "___").sort("outs")

mlb_stuff_plus​

mlb_stuff_plus(pitches: 'pl.DataFrame', *, level: 'str' = 'pitch', return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Score pitches with the bundled Stuff+ (①) run-value model.

Parameters

ParameterTypeDefaultDescription
pitchesDataFrameOutput of sportsdataverse.mlb.mlb_pitch_features.pitch_features (needs velo_z, spin_z, pfx_x_z, pfx_z_z, release_pos_x_z, release_pos_z_z, extension_z).
levelstr'pitch'"pitch" (default) for per-pitch output, or "arsenal" for a per (pitcher, pitch_type) mean.
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

pitcher, pitch_type, stuff_rv_hat, stuff_plus — one row per pitch (level="pitch") or per pitcher-pitchtype (level="arsenal"). Empty input returns a zero-row frame with the documented schema.

col_nametypedescription
pitcherintegerMLB Advanced Media (MLBAM) id for the pitcher.
pitch_typecharacterStatcast pitch-type abbreviation.
stuff_rv_hatdoublePredicted per-pitch run value from the bundled Stuff+ xgboost model (physics + fastball-relative features only).
stuff_plusdoublePlus-scale Stuff+ score, 100 = league average, higher = better (sign-inverted from stuff_rv_hat).

Example

from sportsdataverse.mlb.mlb_pitch_features import pitch_features
from sportsdataverse.mlb.mlb_stuff_plus import mlb_stuff_plus
feats = pitch_features(raw_pitches)
out = mlb_stuff_plus(feats, level="arsenal")
print(out.sort("stuff_plus", descending=True).head())

# Pipeline next step

out.filter(pl.col("pitch_type") == "FF").sort("stuff_plus", descending=True)

mlb_swing_decision​

mlb_swing_decision(start_dt: 'str', end_dt: 'str', *, puller: 'Optional[Callable[..., pl.DataFrame]]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Per player-season swing/take run value + selective-aggression (SEAGER analog).

Pulls pitches via puller(start_dt, end_dt, player_type="batter"), builds the RV(swing)/RV(take) zone x count surfaces (and the league swing-rate table) from the pull itself, then per batter:

  • swing_take_runs = sum of the actual per-pitch delta_run_exp credited to the batter's swing/take decisions (matching Savant's swing/take run-value definition -- the run value of what actually happened on each pitch, not a league-average lookup).
  • selective_agg = sum of rv_chosen - rv_neutral, where rv_neutral = swing_rate * rv_swing + (1 - swing_rate) * rv_take uses the league swing rate for that zone x count cell -- positive means the batter swings at hittable pitches and takes bad ones more than a league-average decision-maker would.
  • chase_rate = swings / pitches seen in the waste/chase zones (zone in {11,12,13,14}).

Parameters

ParameterTypeDefaultDescription
start_dtstrPull start date, YYYY-MM-DD.
end_dtstrPull end date, YYYY-MM-DD.
pullerOptional[Callable[..., DataFrame]]NoneInjectable Statcast search callable -- defaults to sportsdataverse.mlb.mlb_statcast_extra.mlb_statcast_search.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (batter, season): pitches, swing_take_runs, selective_agg, chase_rate, n_swings. Empty pull returns a zero-row frame with the documented schema.

col_nametypedescription
batterintegerMLBAM batter id (join key into Savant's swing/take leaderboard as player_id).
seasonintegerFour-digit season year derived from game_year/game_date.
pitchesintegerTotal pitches seen with a non-null zone/decision in the pulled window.
swing_take_runsdoubleSum of the run value of the batter's actual swing/take decisions (delta_run_exp of the chosen decision at that zone x count).
selective_aggdoubleSEAGER-analog selective-aggression score -- sum of (chosen run value minus the league-neutral-rate run value) per pitch.
chase_ratedoubleShare of pitches in the waste/chase attack zones (11-14) that the batter swung at.
n_swingsintegerCount of pitches the batter swung at.

Example

from sportsdataverse.mlb.mlb_swing_decision import mlb_swing_decision

df = mlb_swing_decision("2024-06-01", "2024-06-21")
print(df.shape)

# Pipeline next step (one line)

df.sort("selective_agg", descending=True).head()

mlb_team_elo​

mlb_team_elo(results: 'pl.DataFrame', *, k: 'float' = 4.0, hfa: 'float' = 24.0, init: 'float' = 1500.0, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

As-of-date iterative Elo run-differential rating.

Games are folded in date order (ties broken by game_id); each team's rating updates only after its game is scored, so the home_rating/away_rating columns are strictly as-of-date (no leakage from later games). home_win_prob_elo uses the standard logistic Elo formula with a home-field-advantage offset.

Parameters

ParameterTypeDefaultDescription
resultsDataFrameGame-level results (game_id, date, home_team_id, away_team_id, home_score, away_score).
kfloat4.0Elo K-factor (rating-update step size).
hfafloat24.0Home-field-advantage Elo-point offset.
initfloat1500.0Initial rating for a team with no prior games.
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

one row per game, in date order. | Column | Type | Description | |---|---|---| | game_id | Utf8 | Game identifier | | date | Date | Game date | | home_team_id | Utf8 | Home team identifier | | away_team_id | Utf8 | Away team identifier | | home_rating | Float64 | Home team's rating before this game | | away_rating | Float64 | Away team's rating before this game | | home_win_prob_elo | Float64 | Elo-implied P(home wins) before this game | | home_rating_post | Float64 | Home team's rating after this game | | away_rating_post | Float64 | Away team's rating after this game |

col_nametypedescription
game_idcharacterGame identifier (statsapi gamePk, stringified).
datedateCalendar date of the game (YYYY-MM-DD).
home_team_idcharacterHome team identifier.
away_team_idcharacterAway team identifier.
home_ratingdoubleHome team's Elo rating before this game (as-of-date).
away_ratingdoubleAway team's Elo rating before this game (as-of-date).
home_win_prob_elodoubleElo-implied P(home team wins) before this game.
home_rating_postdoubleHome team's Elo rating after this game.
away_rating_postdoubleAway team's Elo rating after this game.

Example

from sportsdataverse.mlb.mlb_team_projection import mlb_team_elo
elo = mlb_team_elo(results)

# Pipeline next step (one line)

elo.group_by("home_team_id").agg(pl.col("home_rating_post").last())

mlb_team_projection​

mlb_team_projection(seasons: 'Union[int, List[int], None]' = None, *, results: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Combined pythagenpat + Elo team projection.

Parameters

ParameterTypeDefaultDescription
seasonsUnion[int, List[int], None]NoneReserved for a future network-collector path (currently unused -- pass results directly; see sportsdataverse.mlb.mlb_run_expectancy.mlb_run_expectancy_matrix for the collector pattern this will follow once wired).
resultsOptional[DataFrame]NoneGame-level results (see mlb_pythagenpat_table and mlb_team_elo for the required columns).
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

one row per (season, team). | Column | Type | Description | |---|---|---| | season | Int64 | Season | | team_id | Utf8 | Team identifier | | win_pct | Float64 | Realized win percentage | | pythag_win_pct | Float64 | Pythagenpat expected win percentage | | rating | Float64 | Final (as of the last observed game) Elo rating | | exp_margin | Float64 | Elo-implied expected run margin vs a league-average opponent |

col_nametypedescription
seasonintegerMLB season (4-digit start year).
team_idcharacterTeam identifier (statsapi team id, stringified).
win_pctdoubleRealized win percentage.
pythag_win_pctdoublePythagenpat expected win percentage.
ratingdoubleFinal (as of the last observed game) Elo rating.
exp_margindoubleElo-implied expected run margin vs a league-average opponent.

Example

from sportsdataverse.mlb.mlb_team_projection import mlb_team_projection
projection = mlb_team_projection(results=results)

mlb_win_expectancy​

mlb_win_expectancy(pbp: 'pl.DataFrame', results: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Per-play home win expectancy from the empirical state table.

Parameters

ParameterTypeDefaultDescription
pbpDataFrameParsed mlb_play_by_play frame (see sportsdataverse.mlb.mlb_run_expectancy.pbp_base_out_states).
resultsDataFrameGame-level results (game_id, home_score, away_score).
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

one row per plate appearance, plus one terminal "game over" row per game (at_bat_index = last real PA's index + 1, home_win_exp pinned to the actual final outcome: 1.0 if home won, 0.0 otherwise). Without this anchor, the last real play's own WPA swing (e.g. a walk-off) would never be captured by mlb_win_probability_added's per-game diff, and the game-level WPA sum would not telescope to the exact +-0.5 identity. | Column | Type | Description | |---|---|---| | game_id | Utf8 | Game identifier | | at_bat_index | Int64 | Game-global sequential PA index (last row is a synthetic terminal marker) | | half | Utf8 | "top" or "bottom" (offense side); the terminal row repeats the last real half | | home_win_exp | Float64 | P(home team wins | state before the play); 1.0/0.0 on the terminal row |

col_nametypedescription
game_idcharacterGame identifier (statsapi gamePk, stringified).
at_bat_indexintegerGame-global sequential plate-appearance index.
halfcharacterHalf-inning ("top" or "bottom") -- which side is on offense.
home_win_expdoubleEmpirical P(home team wins | base-out-score-inning state before the play).

Example

from sportsdataverse.mlb.mlb_win_expectancy import mlb_win_expectancy
we = mlb_win_expectancy(pbp, results)

# Pipeline next step (one line)

we.filter(pl.col("game_id") == "716390").sort("at_bat_index")

mlb_win_probability_added​

mlb_win_probability_added(we: 'pl.DataFrame', *, perspective: 'str' = 'home', return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Per-play win-probability-added from a mlb_win_expectancy frame.

wpa_i = home_win_exp_i - home_win_exp_{i-1} within each game (the first play of a game is measured against the neutral 0.5 baseline).

Parameters

ParameterTypeDefaultDescription
weDataFrameOutput of mlb_win_expectancy (needs game_id, at_bat_index, home_win_exp).
perspectivestr'home'"home" (default) returns home-team WPA; any other value (e.g. "away") returns the sign-flipped (away-team) WPA.
return_as_pandasboolFalseReturn pandas.DataFrame instead of polars.

Returns

one row per plate appearance. | Column | Type | Description | |---|---|---| | game_id | Utf8 | Game identifier | | at_bat_index | Int64 | Game-global sequential PA index | | wpa | Float64 | Win-probability added, from perspective |

col_nametypedescription
game_idcharacterGame identifier (statsapi gamePk, stringified).
at_bat_indexintegerGame-global sequential plate-appearance index.
wpadoubleWin-probability added on this play, from the requested perspective.

Example

from sportsdataverse.mlb.mlb_win_expectancy import mlb_win_probability_added
wpa = mlb_win_probability_added(we)

pbp_base_out_states​

pbp_base_out_states(pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Reconstruct pre-play base-out state from statsapi play-by-play.

Within each (game_id, inning, half) half-inning, ordered by the game-global at_bat_index: base_state/outs_start before PA i are the post-occupancy / out-count of PA i-1 (empty/0 at the half's first PA -- occupancy and outs both genuinely reset at every half-inning boundary). runs_on_play is the score delta since the previous PA in the game (over("game_id"), not reset per half-inning -- the score itself carries across the half-inning boundary even though outs/bases do not). runs_rest_of_inning is the suffix-sum of runs_on_play within the half.

Parameters

ParameterTypeDefaultDescription
pbpDataFrameParsed mlb_play_by_play frame (optionally concatenated across games), carrying game_id, about_inning, about_half_inning, about_at_bat_index, count_outs, result_home_score, result_away_score, matchup_post_on_{first,second,third}_id.

Returns

one row per plate appearance. | Column | Type | Description | |---|---|---| | game_id | Utf8 | Game identifier | | inning | Int64 | Inning number | | half | Utf8 | "top" or "bottom" | | at_bat_index | Int64 | Game-global sequential PA index | | base_state | Utf8 | 3-char occupancy before the PA ("1_3" etc.) | | outs_start | Int64 | Outs before the PA (0-2) | | runs_on_play | Int64 | Runs scored on this PA | | runs_rest_of_inning | Int64 | Runs scored from this PA through the half's end | | score_diff | Int64 | home - away score at the start of the PA |

No returns table is published for this function: no capture: its play-by-play input comes from statsapi.mlb.com, which answers HTTP 406 to the datacenter IP the docs are built on, and load_mlb_pbp lacks its game_id / about_* columns.

Example

from sportsdataverse.mlb.mlb_run_expectancy import pbp_base_out_states
states = pbp_base_out_states(pbp)

pearson_corr​

pearson_corr(a: "'np.ndarray'", b: "'np.ndarray'") -> 'float'

Pearson correlation coefficient between two 1-D arrays.

Parameters

ParameterTypeDefaultDescription
andarrayFirst sample array.
bndarraySecond sample array, same length as a.

Returns

Pearson's r. nan if either input has zero variance.

Example

from sportsdataverse.mlb.mlb_run_values import pearson_corr
r = pearson_corr(mine["framing_runs"].to_numpy(), sav["runs_extra_strikes"].to_numpy())

prop_over_prob​

prop_over_prob(line: 'float', expected: 'float') -> 'float'

P(realized count > line) under a Poisson(expected) model.

1 - poisson.cdf(floor(line), expected).

Parameters

ParameterTypeDefaultDescription
linefloatThe prop betting line (e.g. 8.5 runs).
expectedfloatThe Poisson mean (expected runs/strikeouts/etc.).

Returns

P(over), in [0, 1].

Example

from sportsdataverse.mlb.mlb_prop_projection import prop_over_prob
prop_over_prob(3.5, 4.5)

spearman_corr​

spearman_corr(a: 'np.ndarray', b: 'np.ndarray') -> 'float'

Spearman rank correlation between two arrays.

Parameters

ParameterTypeDefaultDescription
andarrayFirst array of values.
bndarraySecond array of values (same length as a).

Returns

The Spearman rank correlation coefficient.

Example

import numpy as np
from sportsdataverse._common.metrics import spearman_corr
spearman_corr(np.array([1, 2, 3]), np.array([3, 1, 2]))