Skip to main content
Version: 0.1.5

NBA — additional Python functions — IDs and crosswalks

nba_player_crosswalk​

nba_player_crosswalk(season: 'Optional[int]' = None, min_confidence: 'float' = 0.92, *, return_as_pandas: 'bool' = False, strict: 'bool' = False, **kwargs: 'Any') -> "Union[pl.DataFrame, 'pd.DataFrame']"

Build the NBA cross-source player crosswalk (ESPN / NBA Stats / Fox).

One row per ESPN athlete per team. match_method / match_confidence describe the Stats API match (normalized exact name, then Jaro-Winkler with jersey and DOB tiebreaks); Fox contributes fox_athlete_id only.

Parameters

ParameterTypeDefaultDescription
seasonOptional[int]NoneSeason year per hoopR convention. Defaults to the most recent NBA season.
min_confidencefloat0.92Jaro-Winkler floor for fuzzy matches (R default 0.92).
return_as_pandasboolFalseReturn pandas instead of polars.
strictboolFalseRaise on the first failed per-team ESPN or Fox roster fetch (a 404 is still skipped) instead of skipping isolated failures. Default False matches the R producers; a provider whose every item failed raises either way. An item the host answered -- including a 404 -- counts as answered.

Returns

pl.DataFrame (or pandas), one row per ESPN athlete, 21 columns.

col_nametypedescription
seasonintegerSeason year.
espn_team_idintegerESPN team id (canonical key).
team_abbreviationcharacterShort team abbreviation (e.g. 'LAS').
player_namecharacterPlayer name.
espn_athlete_idcharacterESPN athlete id.
espn_full_namecharacterESPN full name.
espn_jerseycharacterESPN jersey number.
espn_positioncharacterESPN position abbreviation.
nba_player_idcharacterNBA Stats API (stats.nba.com) player id as a string, matched to the ESPN athlete within the same team by normalized exact name, then Jaro-Winkler fuzzy name match (min_confidence, default 0.92) with jersey and birth-date tiebreaks; null when the athlete had no Stats match.
nba_player_namecharacterPlayer name from the NBA Stats commonteamroster row matched to the ESPN athlete; null when the athlete had no Stats match.
nba_jersey_numcharacterJersey number as a string from the NBA Stats commonteamroster row matched to the ESPN athlete; null when the athlete had no Stats match.
nba_positioncharacterPosition from the NBA Stats commonteamroster row matched to the ESPN athlete; null when the athlete had no Stats match.
fox_athlete_idcharacterFox athlete id (NA if unmatched).
fox_playercharacterFox player name (NA if unmatched).
fox_jerseycharacterFox jersey number (NA if unmatched).
fox_position_groupcharacterFox position group label (NA if unmatched).
yahoo_player_idcharacterYahoo player id (NA placeholder).
yahoo_player_namecharacterYahoo player name (NA placeholder).
match_methodcharacterCombination of matched sources, e.g. "fox+bart" / "fox_only" / "bart_only" / "espn_only".
match_confidencedoubleJaro-Winkler score or 1 for exact (NA if none).
match_keyscharacterNA (reserved for future use).

Example

from sportsdataverse.nba import nba_player_crosswalk
df = nba_player_crosswalk(season=2026)
print(df["match_method"].value_counts())

# Tighten the fuzzy floor

strict = nba_player_crosswalk(season=2026, min_confidence=0.97)

# Pipeline next step (one line)

df.filter(pl.col("match_method") == "fuzzy_jw").head()

nba_schedule_crosswalk​

nba_schedule_crosswalk(season: 'Optional[int]' = None, *, stats_games: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False, strict: 'bool' = False, **kwargs: 'Any') -> "Union[pl.DataFrame, 'pd.DataFrame']"

Build the NBA cross-source schedule crosswalk (ESPN / NBA Stats).

One row per game. Both sides reduce to the Eastern-Time game date before joining on (game_date, home_espn_team_id, away_espn_team_id). The Stats CDN serves the current season only, so the live builder is effectively current-season.

Parameters

ParameterTypeDefaultDescription
seasonOptional[int]NoneSeason year per hoopR convention. Defaults to the most recent NBA season.
stats_gamesOptional[DataFrame]NonePre-fetched Stats schedule frame; None fetches live.
return_as_pandasboolFalseReturn pandas instead of polars.
strictboolFalseRaise on the first failed per-date ESPN scoreboard fetch (a 404 is still skipped) instead of skipping isolated failures. Default False matches the R producers; a provider whose every item failed raises either way. An item the host answered -- including a 404 -- counts as answered.

Returns

pl.DataFrame (or pandas) with SCHEDULE_COLUMNS.

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba import nba_schedule_crosswalk
df = nba_schedule_crosswalk(season=2026)
print(df["match_method"].value_counts())

# Pipeline next step (one line)

df.filter(pl.col("match_method") == "both").select("espn_game_id", "nba_game_id").head()

nba_team_crosswalk​

nba_team_crosswalk(season: 'Optional[int]' = None, *, stats: 'Optional[pl.DataFrame]' = None, fox: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False, **kwargs: 'Any') -> "Union[pl.DataFrame, 'pd.DataFrame']"

Build the NBA cross-source team crosswalk (ESPN / NBA Stats / Fox).

One row per ESPN team, keyed on espn_team_id. ESPN and Stats team endpoints are current-season snapshots, so season is a stamp; historical relocations are not back-modelled.

Parameters

ParameterTypeDefaultDescription
seasonOptional[int]NoneSeason year per hoopR convention (2026 = 2025-26). Defaults to the most recent NBA season.
statsOptional[DataFrame]NonePre-fetched Stats team directory (espn_team_id + nba_team_*). None derives it from nba_stats_leaguestandingsv3 joined to ESPN on the normalized team nickname.
foxOptional[DataFrame]NonePre-fetched Fox directory. None fetches live.
return_as_pandasboolFalseReturn pandas instead of polars.

Returns

pl.DataFrame (or pandas), one row per ESPN team, with TEAM_COLUMNS.

col_nametypedescription
seasonintegerSeason year.
espn_team_idintegerESPN team id (canonical key).
espn_abbreviationcharacterESPN abbreviation.
espn_display_namecharacterESPN display name (school + mascot).
espn_short_namecharacterESPN short name.
espn_locationcharacterESPN school/location only.
espn_mascotcharacterESPN mascot/nickname.
nba_team_idcharacterNBA Stats API (stats.nba.com) team id as a string, attached to the ESPN team row on espn_team_id after the Stats team nickname is matched to ESPN's short_name; null when no Stats team matched the ESPN team.
nba_team_abbreviationcharacterNBA Stats team tricode, taken from nba_stats_leaguegamelog's team_abbreviation because leaguestandingsv3 publishes none; null when no Stats team matched the ESPN team.
nba_team_namecharacterFull NBA Stats team name built as team_city plus team_name (city then nickname) from nba_stats_leaguestandingsv3; null when no Stats team matched the ESPN team.
nba_team_citycharacterTeam city (team_city) from nba_stats_leaguestandingsv3; null when no Stats team matched the ESPN team.
nba_team_slugcharacterURL slug for the team (team_slug) from nba_stats_leaguestandingsv3; null when no Stats team matched the ESPN team.
nba_conferencecharacterTeam's conference as NBA Stats labels it (conference) from nba_stats_leaguestandingsv3; null when no Stats team matched the ESPN team.
nba_divisioncharacterTeam's division as NBA Stats labels it (division) from nba_stats_leaguestandingsv3; null when no Stats team matched the ESPN team.
fox_team_idcharacterFox Bifrost team id (NA if unmatched).
fox_team_namecharacterFox team name (NA if unmatched).
yahoo_team_idcharacterYahoo team id (NA placeholder).
yahoo_team_abbreviationcharacterYahoo abbreviation (NA placeholder).
yahoo_team_namecharacterYahoo team name (NA placeholder).
match_methodcharacterCombination of matched sources, e.g. "fox+bart" / "fox_only" / "bart_only" / "espn_only".
match_confidencedoubleJaro-Winkler score or 1 for exact (NA if none).

Example

from sportsdataverse.nba import nba_team_crosswalk
df = nba_team_crosswalk(season=2026)
print(df.shape)

# Offline with a pre-fetched Stats frame

df = nba_team_crosswalk(season=2026, stats=my_stats, fox=my_fox)

# Pipeline next step (one line)

df.select("espn_team_id", "nba_team_id", "match_method").head()