Skip to main content

WNBA — additional Python functions — Other

load_wnba_stats_leaguedash​

load_wnba_stats_leaguedash(family: 'str', seasons, return_as_pandas: 'bool' = False) -> 'pl.DataFrame'

Load one asset family of the wnba_stats_leaguedash release.

wnba_stats_leaguedash is a parameter cube: one asset per (family, season) pair rather than one per season, so a family must be named. The valid families are exported as WNBA_STATS_LEAGUEDASH_FAMILIES -- import that tuple to discover them rather than passing a bare string; an unknown family raises ValueError listing every valid value. This is the non-deprecated way to reach the cube; the four load_wnba_stats_* shims below only reconstruct retired tags' stacked shapes from it.

Column sets are family-specific (a lineups_* frame keys on group_id, a player_* frame on player_id), so this loader documents no fixed returns table. player_id / team_id are Int64 in every family and season, so cross-family joins need no dtype reconciliation.

Parameters

ParameterTypeDefaultDescription
familystrAsset family, e.g. "player_stats_advanced". Must be one of WNBA_STATS_LEAGUEDASH_FAMILIES.
seasonsint | Iterable[int]Season, or iterable of seasons, to load. WNBA seasons are single calendar years. 1997 is the earliest season on the tag. A requested season the family does not publish is warned about and skipped, not an error.
return_as_pandasboolFalseIf True, returns a pandas dataframe. If False, returns a polars dataframe.

Returns

Polars dataframe with one row per player / team / lineup per requested season for the requested family; an empty frame when no requested season is published.

Example

from sportsdataverse.wnba import load_wnba_stats_leaguedash
adv = load_wnba_stats_leaguedash("player_stats_advanced", seasons=2025)
print(adv.shape)

# Discover the valid families

from sportsdataverse.wnba import WNBA_STATS_LEAGUEDASH_FAMILIES
print(WNBA_STATS_LEAGUEDASH_FAMILIES)

# Multi-season, pandas round-trip

team_pd = load_wnba_stats_leaguedash(
"team_stats_base", seasons=range(2020, 2026), return_as_pandas=True
)

# Pipeline next step (best net rating in 2025)

import polars as pl
load_wnba_stats_leaguedash("team_stats_advanced", seasons=2025).sort(
"net_rating", descending=True
).head()

load_wnba_stats_lineups​

load_wnba_stats_lineups(seasons, return_as_pandas: 'bool' = False) -> 'pl.DataFrame'

Load season-level WNBA 5-man lineup statistics (deprecated).

Parameters

ParameterTypeDefaultDescription
seasonsan int or iterable of seasons.
return_as_pandasboolFalsereturn a pandas DataFrame instead of polars.

Returns

A polars (or pandas) DataFrame, one row per lineup-season-measure_type, stacked from the wnba_stats_leaguedash cube's lineups_{base, advanced} assets filtered to group_quantity == 5 — matching the old wnba_stats_lineups tag's 5-man-only, Base+Advanced-only coverage. Call the cube's lineups_* assets directly (unfiltered) for 2/3/4-man lineups or the other 4 measure types.

Example

from sportsdataverse.wnba import load_wnba_stats_lineups
df = load_wnba_stats_lineups(seasons=2026)
print(df.shape)

load_wnba_stats_player_season_stats​

load_wnba_stats_player_season_stats(seasons, return_as_pandas: 'bool' = False) -> 'pl.DataFrame'

Load season-level WNBA player statistics (deprecated).

Parameters

ParameterTypeDefaultDescription
seasonsan int or iterable of seasons.
return_as_pandasboolFalsereturn a pandas DataFrame instead of polars.

Returns

A polars (or pandas) DataFrame, one row per player-season-measure_type, stacked from the wnba_stats_leaguedash cube's player_stats_* assets (Base/Advanced/Misc/Scoring/Usage/Defense — matches the old wnba_stats_player_season_stats tag's coverage; player-level Opponent/Four Factors are empty upstream and were never populated by either version).

Example

from sportsdataverse.wnba import load_wnba_stats_player_season_stats
df = load_wnba_stats_player_season_stats(seasons=2026)
print(df.shape)

# Pipeline next step (Advanced-only rows)

import polars as pl
adv = df.filter(pl.col("measure_type") == "Advanced")

load_wnba_stats_standings​

load_wnba_stats_standings(seasons, return_as_pandas: 'bool' = False) -> 'pl.DataFrame'

Load season-level WNBA standings (deprecated).

Parameters

ParameterTypeDefaultDescription
seasonsan int or iterable of seasons.
return_as_pandasboolFalsereturn a pandas DataFrame instead of polars.

Returns

A polars (or pandas) DataFrame, one row per team-season, read from the wnba_stats_leaguedash cube's standings asset -- the same underlying leaguestandingsv3 endpoint/params as the old wnba_stats_standings tag, so this is close to a pure passthrough.

Example

from sportsdataverse.wnba import load_wnba_stats_standings
df = load_wnba_stats_standings(seasons=2026)
print(df.shape)

load_wnba_stats_team_season_stats​

load_wnba_stats_team_season_stats(seasons, return_as_pandas: 'bool' = False) -> 'pl.DataFrame'

Load season-level WNBA team statistics (deprecated).

Parameters

ParameterTypeDefaultDescription
seasonsan int or iterable of seasons.
return_as_pandasboolFalsereturn a pandas DataFrame instead of polars.

Returns

A polars (or pandas) DataFrame, one row per team-season-measure_type, stacked from the wnba_stats_leaguedash cube's team_stats_* assets (Base/Advanced/Misc/Scoring/Defense/ Opponent — matches the old wnba_stats_team_season_stats tag's coverage; team-level Usage/Four Factors are empty upstream).

Example

from sportsdataverse.wnba import load_wnba_stats_team_season_stats
df = load_wnba_stats_team_season_stats(seasons=2026)
print(df.shape)

most_recent_wnba_season​

most_recent_wnba_season()

most_recent_wnba_season - return the most recent (likely-completed) WNBA season year.

Returns the current calendar year if it's May or later (the WNBA regular season has tipped off), otherwise the previous calendar year.

Returns

Year (e.g. 2024) suitable for passing as a season argument to schedule / loader functions.

Example

from sportsdataverse.wnba import most_recent_wnba_season, espn_wnba_calendar
season = most_recent_wnba_season()
cal = espn_wnba_calendar(season=season)
print(season, cal.height)

build_athlete_identity_lookup​

build_athlete_identity_lookup(rosters: 'dict[int | str, dict]') -> 'dict[str, dict[str, Any]]'

R build_athlete_identity_lookup: athlete_id -> identity from team rosters.

Parameters

ParameterTypeDefaultDescription
rostersdict[int | str, dict]Mapping of team_id -> that team's raw roster payload (wbb/team_rosters/json/{season}/{team_id}.json). NOTE: R walks raw$athletes directly here (no position-bucket unwrap, unlike the rosters dataset itself).

Returns

athlete_id (str) -> identity fields for helper_wbb_player_season_stats.

build_wnba_season_wp​

build_wnba_season_wp(season: 'int', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

A WNBA season's play-by-play with win-probability columns joined in.

Loads the season's play-by-play, schedule, and team boxscores, builds a leakage-free weekly as-of pregame anchor per game from the WNBA ratings engine (league_id="10"), scores every play through the bundled in-game win-probability artifact, and returns the full load_wnba_pbp frame with pregame_home_prob + home_win_prob appended -- the enrich-in-place shape that overwrites the season's play_by_play_<season>.parquet release asset.

Parameters

ParameterTypeDefaultDescription
seasonintSeason year (e.g. 2024); bounded by load_wnba_pbp release availability.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

The season's load_wnba_pbp frame (every column preserved) with the two WP columns pregame_home_prob + home_win_prob appended (both Float64), sorted by game_id then game_play_number.

Example

from sportsdataverse.wnba import build_wnba_season_wp
wp = build_wnba_season_wp(2024)
wp.select("game_id", "game_play_number", "home_win_prob").head()

# Pandas output

wp_pd = build_wnba_season_wp(2024, return_as_pandas=True)

espn_wnba_teams​

espn_wnba_teams(return_as_pandas=False, **kwargs) -> 'pl.DataFrame'

espn_wnba_teams - look up WNBA teams

Parameters

ParameterTypeDefaultDescription
return_as_pandasboolFalseIf True, returns a pandas dataframe. If False, returns a polars dataframe.

Returns

Polars dataframe containing teams for the requested league. This function caches by default, so if you want to refresh the data, use the command sportsdataverse.wnba.espn_wnba_teams.clear_cache().

col_nametypedescription
team_abbreviationcharacterShort team abbreviation (e.g. 'LAS').
team_alternate_colorcharacterTeam alternate color (hex without leading '#').
team_colorcharacterTeam primary color (hex without leading '#').
team_display_namecharacterFull team display name.
team_idcharacterUnique team identifier.
team_is_activelogicalTRUE if the team is currently active.
team_is_all_starlogicalTRUE if the row represents an All-Star team.
team_locationcharacterTeam city or location string.
team_logosintegerTeam logo metadata.
team_namecharacterFull team display name (e.g. 'Las Vegas Aces').
team_nicknamecharacterTeam nickname.
team_short_display_namecharacterShort team display name (e.g. 'Aces').
team_slugcharacterURL-safe team identifier (e.g. 'lasvegas-aces' / 'aces').
team_uidcharacterESPN universal team identifier (UID format 's:40~l:...~t:...').

Example

from sportsdataverse.wnba import espn_wnba_teams
teams = espn_wnba_teams()
print(teams.shape)
teams.select(["team_id", "team_abbreviation", "team_display_name"]).head()

# Find Las Vegas Aces (team_id 17)

teams.filter(__import__("polars").col("team_id") == "17").to_dicts()

# Refresh the cache (the call is ``lru_cache``'d)

espn_wnba_teams.cache_clear() # cached at function-level
teams_pd = espn_wnba_teams(return_as_pandas=True)

make_prob_by_context​

make_prob_by_context(ptshots: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "'dict[str, Union[pl.DataFrame, pd.DataFrame]]'"

Marginal FG% tables by defender distance and by shot clock.

The public API exposes defender-distance and shot-clock only as aggregate bucket tables (playerdashptshots), not per-shot fields, so this aggregates Σfgm/Σfga across players within each bucket.

Parameters

ParameterTypeDefaultDescription
ptshotsDataFrameThe stacked playerdashptshots fixture — one frame with a result_set tag (ClosestDefenderShooting / ShotClockShooting) plus bucket, fga, fgm.
return_as_pandasboolFalseReturn pandas DataFrames instead of polars.

Returns

{"defender": frame, "shot_clock": frame} each with rows per bucket (bucket, fga, fgm, fg_pct). Missing result sets return the zero-row schema.

Example

from sportsdataverse.nba.nba_shot_value import make_prob_by_context
tables = make_prob_by_context(ptshots)
tables["defender"].sort("fg_pct")

make_prob_joint​

make_prob_joint(defender: 'pl.DataFrame', shot_clock: 'pl.DataFrame', overall_fg_pct: 'float', *, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Independence-combined defender x shot-clock make probability.

Combines the two marginal FG% tables under a conditional-independence assumption via odds multipliers: odds(p) = p/(1-p); odds_joint = odds_overall * (odds_def/odds_overall) * (odds_clock/odds_overall); joint = odds_joint/(1+odds_joint). This assumes defender distance and shot-clock effects are independent given the league baseline — a simplification (a late clock correlates with tighter defense), documented here so callers weigh it.

Parameters

ParameterTypeDefaultDescription
defenderDataFrameThe "defender" marginal table from make_prob_by_context (bucket, fg_pct).
shot_clockDataFrameThe "shot_clock" marginal table (bucket, fg_pct).
overall_fg_pctfloatThe league overall FG% baseline.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (close_def_dist_range, shot_clock_range): close_def_dist_range:Utf8, shot_clock_range:Utf8, joint_fg_pct:Float64. Empty inputs return the zero-row schema.

Example

from sportsdataverse.nba.nba_shot_value import make_prob_by_context, make_prob_joint
t = make_prob_by_context(ptshots)
joint = make_prob_joint(t["defender"], t["shot_clock"], 0.47)

score_shot_xpoints​

score_shot_xpoints(shots: 'pl.DataFrame', league_avgs: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Score each shot with expected points from the league-average baseline.

Joins the per-shot frame to the zone baseline (falling back to the within-shot_zone_range mean when a zone triple is unmatched) and adds shot_value (3 for a 3PT shot else 2), xpoints = base_fg_pct * shot_value, and actual_points = shot_made_flag * shot_value.

Parameters

ParameterTypeDefaultDescription
shotsDataFramePer-shot Shot_Chart_Detail frame (needs shot_type + the three zone keys + shot_made_flag).
league_avgsDataFrameThe LeagueAverages frame (see xpoints_baseline).
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

The input shots plus shot_value:Int64, base_fg_pct:Float64, xpoints:Float64, actual_points:Float64. Empty input returns the augmented schema with zero rows.

Example

from sportsdataverse.nba.nba_shot_value import score_shot_xpoints
scored = score_shot_xpoints(shots, league_avgs)

# Pipeline next step (one line)

scored.group_by("player_id").agg(pl.col("xpoints").sum())

scoreboard_event_parsing​

scoreboard_event_parsing(event)

No description available.

Parameters

ParameterTypeDefaultDescription
event

shooter_talent​

shooter_talent(scored_shots: 'pl.DataFrame', *, league_id: 'str' = '00', min_attempts: 'int' = 50, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Regressed shooter true-talent: make%-above-expected, shrunk to the mean.

Aggregates score_shot_xpoints output per shooter and regresses the raw over-expected rate toward zero by n/(n+k) (k = get_shrinkage_k(league_id), fitted split-half). As-of leakage boundary: to score a shooter's talent for shots after date D, pass only that shooter's shots before D -- this function does not enforce the cut itself.

Parameters

ParameterTypeDefaultDescription
scored_shotsDataFramescore_shot_xpoints output (needs player_id, shot_made_flag, base_fg_pct, xpoints, actual_points).
league_idstr'00'"00" NBA, "10" WNBA, "20" G-League.
min_attemptsint50Drop shooters with fewer attempts (unstable estimate).
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per player_id: player_id:Int64, n_att:Int64, actual_makes:Int64, exp_makes:Float64, points_above_expected:Float64, raw_above_pct:Float64, talent_pct:Float64. Empty input returns the zero-row schema.

Example

from sportsdataverse.nba.nba_shot_value import score_shot_xpoints, shooter_talent
talent = shooter_talent(score_shot_xpoints(shots, league_avgs))

# Pipeline next step (one line)

talent.sort("talent_pct", descending=True).head(15)

shot_selection_quality​

shot_selection_quality(scored_shots: 'pl.DataFrame', *, min_attempts: 'int' = 50, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Player shot-selection quality: mean expected value vs the league mean.

xev_per_shot is a player's mean xpoints (the value of the LOOKS they take, independent of makes); selection_quality is that minus the league-wide mean xpoints over the same frame -- a rim-and-three diet scores positive, a mid-range diet negative.

Parameters

ParameterTypeDefaultDescription
scored_shotsDataFramescore_shot_xpoints output (needs player_id, xpoints).
min_attemptsint50Drop players with fewer attempts.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per player_id: player_id:Int64, n_att:Int64, xev_per_shot:Float64, league_xev_per_shot:Float64, selection_quality:Float64. Empty input returns the zero-row schema.

Example

from sportsdataverse.nba.nba_shot_value import score_shot_xpoints, shot_selection_quality
sel = shot_selection_quality(score_shot_xpoints(shots, league_avgs))

# Pipeline next step (one line)

sel.sort("selection_quality", descending=True).head(15)

zone_value_map​

zone_value_map(scored_shots: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Per-player per-zone value map: points and expected points per shot.

Collapses shot_zone_basic to a canonical zone via ZONE_COLLAPSE (the two corner-3 zones merge) and aggregates realized vs expected points per shot in each zone.

Parameters

ParameterTypeDefaultDescription
scored_shotsDataFramescore_shot_xpoints output (needs player_id, shot_zone_basic, shot_made_flag, actual_points, xpoints).
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (player_id, zone): player_id:Int64, zone:Utf8, att:Int64, makes:Int64, pts:Float64, pps:Float64, xpps:Float64, pps_above_expected:Float64 (pps = points per shot, xpps = expected). Empty input returns the zero-row schema.

Example

from sportsdataverse.nba.nba_shot_value import score_shot_xpoints, zone_value_map
zmap = zone_value_map(score_shot_xpoints(shots, league_avgs))

# Pipeline next step (one line)

zmap.filter(pl.col("zone") == "corner_3").sort("pps_above_expected", descending=True)