Skip to main content

MBB — additional Python functions — NCAA (stats.ncaa.org)

ncaa_espn_team_crosswalk​

ncaa_espn_team_crosswalk(league: 'str' = 'mbb', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Season-keyed stats.ncaa.org -> ESPN team-id crosswalk.

One row per (season, ncaa_team_id). Teams that could not be resolved to an ESPN team are kept with a null espn_team_id and match_method="unmatched" -- never dropped -- so the row count always equals ncaa_{league}_team_ids().

Parameters

ParameterTypeDefaultDescription
leaguestr'mbb'"mbb" (men's, 2009-10 onward) or "wbb" (women's).
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

DataFrame with columns season (str, "YYYY-YY"), ncaa_team_id (Int64 -- the season-specific stats.ncaa.org id), ncaa_team / ncaa_conference (str), conference_id (str, nullable -- the SDV group id, e.g. "mbb:big-ten"), espn_team_id (str, nullable -- ESPN ids are strings throughout sdv-py), espn_display_name / espn_location / espn_mascot / espn_abbreviation / espn_conference_name / espn_conference_id (str, nullable), and match_method (str -- "exact", "dict", "alias" or "unmatched"). The three conference columns and ncaa_conference are per season: Maryland is ACC in 2013-14 and Big Ten in 2014-15.

col_nametypedescription
seasoncharacterSeason year.
ncaa_team_idintegerstats.ncaa.org team id (Int64) for that season; stats.ncaa.org issues a new id every season, so the same school has a different id on each season row.
ncaa_teamcharacterSchool name as stats.ncaa.org writes it, in AP-style abbreviations (e.g. 'Alabama St.', 'A&M-Corpus Christi').
ncaa_conferencecharacterstats.ncaa.org label of the team's conference that season (e.g. 'SEC', 'MWC'), taken from the groups tables' NCAA aliases for the season's conference_id. A conference with no NCAA alias gets its SDV abbreviation (men's Great West -> 'GWC'); a team the groups table has no row for keeps the bundled stats.ncaa.org team-list label. The label style can differ by league ('MWC' in men's, 'Mountain West' in women's through 2022-23).
espn_team_idcharacterESPN team id (canonical key).
espn_display_namecharacterESPN display name (school + mascot).
espn_locationcharacterESPN school/location only.
espn_mascotcharacterESPN mascot/nickname.
espn_abbreviationcharacterESPN abbreviation.
espn_conference_namecharacterConference name for that season as the {mbb,wbb}_group_seasons table records it (e.g. 'Colonial Athletic Association' through 2022-23, 'Coastal Athletic Association' after). Null on the same rows as conference_id.
espn_conference_idcharacterESPN conference (group) id for that season as a string (e.g. '23' for the Southeastern Conference). ESPN group ids are sport-scoped (Summit League is 49 in men's, 47 in women's) and can change (men's Summit League was 15 before 2008). Null on the same rows as conference_id.
match_methodcharacterCombination of matched sources, e.g. "fox+bart" / "fox_only" / "bart_only" / "espn_only".
conference_idcharacterSportsDataverse conference (group) id for that season, prefixed by the league (e.g. 'mbb:big-ten'), from the {mbb,wbb}_team_group_seasons release table joined on espn_team_id and the season's ending year. Stable across renames and shared with the {mbb,wbb}_groups tables; null only when that table has no row for the team that season.

Example

from sportsdataverse.mbb import ncaa_espn_team_crosswalk
df = ncaa_espn_team_crosswalk()
print(df.shape)

# Women's crosswalk as pandas

wdf = ncaa_espn_team_crosswalk(league="wbb", return_as_pandas=True)

# Pipeline next step (one line)

df.filter(pl.col("season") == "2025-26").select("ncaa_team_id", "espn_team_id")

ncaa_mbb_box_scores​

ncaa_mbb_box_scores(game_ids: 'Union[str, int, Iterable[Union[str, int]]]', *, multi_games: 'bool' = False, fetcher: 'Optional[_SupportsFetchIndividualStats]' = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Box scores for one or more NCAA games (bigballR get_box_scores port).

Multi-game driver over parse_ncaa_bb_box (bigballR/R/all_functions.R:3603-3678): drops null ids, isolates per-game errors (failed ids are reported and skipped), binds rows, and optionally aggregates across games.

Parameters

ParameterTypeDefaultDescription
game_idsUnion[str, int, Iterable[Union[str, int]]]One id or an iterable of NCAA contest ids.
multi_gamesboolFalseWhen True, aggregate one row per (player, clean_name, team) -- counters summed, g = games played, rates recomputed from the sums (R's multi.games; grouping adapted per module docstring).
fetcherOptional[_SupportsFetchIndividualStats]NoneOptional injected fetcher exposing fetch_game_individual_stats (for offline replay/tests). Defaults to a fresh NcaaFetcher.with_browser() context per call -- stats.ncaa.org sits behind an Akamai challenge that the plain transport cannot clear.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

polars.DataFrame (or pandas with return_as_pandas=True): per-game rows in the parse_ncaa_bb_box contract, or the aggregated multi_games contract.

Example

from sportsdataverse.mbb.mbb_ncaa_box_stats import ncaa_mbb_box_scores
df = ncaa_mbb_box_scores(["6470186", "6479639"])
print(df.shape)

# Season aggregate for a scraped id list

agg = ncaa_mbb_box_scores(ids, multi_games=True)

# Offline with an injected fetcher

df = ncaa_mbb_box_scores("6470186", fetcher=my_fetcher)

ncaa_mbb_date_games​

ncaa_mbb_date_games(date: 'Optional[str]' = None, *, conference: 'str' = 'All', conference_id: 'Optional[int]' = None, fetcher: "'Optional[NcaaFetcher]'" = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Discover every NCAA MBB game played on a date (bigballR get_date_games).

Fetches stats.ncaa.org/season_divisions/{sid}/scoreboards for the date's season and returns one row per game with the /contests/{id} game id needed by the play-by-play / box-score scrapers. Port of bigballR get_date_games (all_functions.R:1119-1427).

Parameters

ParameterTypeDefaultDescription
dateOptional[str]None"MM/DD/YYYY". Defaults to yesterday (R default).
conferencestr'All'Conference name filter (e.g. "ACC", "Big Ten", "Metro" / "MAAC"); case/punctuation-insensitive, and both the current stats.ncaa.org label and bigballR's abbreviation work. Default "All" (every conference). Unknown names raise.
conference_idOptional[int]NoneExplicit stats.ncaa.org conference id; overrides conference when given (R's conference.ID).
fetcherOptional[NcaaFetcher]NoneInjectable ~sportsdataverse.mbb.mbb_ncaa_fetch .NcaaFetcher (tests pass an offline fake). None uses NcaaFetcher.with_browser() — the page is JS-rendered behind Akamai bm-verify, so the browser transport is the live default.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per game with columns date, start_time, home, away, box_id, game_id, home_score, away_score, attendance, neutral_site, home_wins, home_losses, away_wins, away_losses (SCOREBOARD_SCHEMA). Scores stay Utf8 — they hold "Canceled" / "Ppd" for unplayed games; game_id is null for games without a box score.

Example

from sportsdataverse.mbb.mbb_ncaa_scoreboard import ncaa_mbb_date_games
games = ncaa_mbb_date_games("11/11/2025")
print(games.shape)

# Useful parameter combination

acc_pd = ncaa_mbb_date_games("02/01/2025", conference="ACC",
return_as_pandas=True)

# Pipeline next step (one line)

games.filter(pl.col("game_id").is_not_null())["game_id"].to_list()

ncaa_mbb_game_pbp​

ncaa_mbb_game_pbp(game_id: 'object', *, fetcher: 'Optional[_SupportsFetchGamePbp]' = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, Any]'"

Scrape one MBB game's play-by-play (bigballR scrape_game).

Fetches stats.ncaa.org/contests/{game_id}/play_by_play and parses it through parse_ncaa_bb_game_pbp with the MBB period model (2, 1200, 300).

Parameters

ParameterTypeDefaultDescription
game_idobjectNCAA contest id (e.g. "6470186").
fetcherOptional[_SupportsFetchGamePbp]NoneOptional injected fetcher exposing fetch_game_pbp (for tests/offline use). Defaults to a fresh NcaaFetcher.with_browser() context per call.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

The 35-column play-by-play frame (zero rows when the game is not found).

Example

from sportsdataverse.mbb.mbb_ncaa_game_pbp import ncaa_mbb_game_pbp
df = ncaa_mbb_game_pbp("6470186")
print(df.shape)

# Offline with an injected fetcher

df = ncaa_mbb_game_pbp("6470186", fetcher=my_fetcher)

# Pipeline next step (one line)

df.filter(pl.col("event_type") == "Three Point Jumper").head()

ncaa_mbb_join_pbp_shots​

ncaa_mbb_join_pbp_shots(pbp: 'pl.DataFrame', shots: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, Any]'"

Attach shot-chart coordinates to play-by-play rows (bigballR

join_pbp_shots, get_shot_locations.R:93-135).

FG attempts (shot_value 2/3) are matched to chart shots on (game_id, game_seconds, event_result == shot_result, shot_no) where shot_no is the within-second same-result sequence number on BOTH sides — free throws and non-shot rows are deliberately excluded from matching (the chart plots FGs only) and pass through NA-filled. Row count and per-game row order are preserved; the explicit shot_dist carry-through is fork-skew fix #11 (wbigballR's unexported copy drops it). Works identically for the WBB extension — feed it quarter-model pbp + shots frames.

Parameters

ParameterTypeDefaultDescription
pbpDataFrameThe 35-column snake_case pbp contract frame (see ~sportsdataverse.mbb.mbb_ncaa_game_pbp.PBP_SCHEMA).
shotsDataFrameA SHOTS_SCHEMA frame (from parse_ncaa_bb_shots / ncaa_mbb_shot_locations) covering exactly the same game ids.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

The input pbp frame + team, player, x, y, shot_dist (null on non-FG rows and unmatched FG rows), sorted by (game_id, original per-game row order) exactly as R's arrange(row, .by_group = TRUE).

Example

from sportsdataverse.mbb.mbb_ncaa_game_pbp import ncaa_mbb_play_by_play
from sportsdataverse.mbb.mbb_ncaa_shots import (
ncaa_mbb_join_pbp_shots,
ncaa_mbb_shot_locations,
)
pbp = ncaa_mbb_play_by_play(["6470186"])
shots = ncaa_mbb_shot_locations(["6470186"])
joined = ncaa_mbb_join_pbp_shots(pbp, shots)

# Pipeline next step (one line)

joined.filter(pl.col("x").is_not_null()).head()

ncaa_mbb_lineups​

ncaa_mbb_lineups(pbp: 'pl.DataFrame', *, include_transition: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into per-lineup stats.

Port of bigballR get_lineups (all_functions.R:1945-2521). Rows with any missing on-court player and substitution rows are dropped, each row's home/away five are byte-sorted so a lineup always occupies the same columns, and the home + away passes are combined per (p1..p5, team). Ratios are derived from the summed counters, rounded to 3 decimals, and NA/Inf are zeroed exactly where R does it.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (parse_ncaa_bb_game_pbp output). May span multiple games.
include_transitionboolFalseWhen True, append the trans/half split surface plus o_trans_pct/d_trans_pct (213 columns total).
fix_tip_inboolTrueWhen True (default), rim stats count the scrape engine's real "Tip In" vocabulary. When False, reproduce R's literal "Tip-In" test (all_functions.R:2012) — tip-ins silently excluded — for oracle parity.
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with one row per lineup+team — 75 columns (LINEUPS_COLUMNS) or 213 with include_transition=True (LINEUPS_TRANSITION_COLUMNS), rows sorted by p1..p5, team. Empty input yields an empty frame with the documented schema.

Example

from sportsdataverse.mbb.mbb_ncaa_lineups import ncaa_mbb_lineups
lineups = ncaa_mbb_lineups(pbp)
print(lineups.shape)

# Transition/half-court splits, pandas out

df_pd = ncaa_mbb_lineups(pbp, include_transition=True, return_as_pandas=True)

# Pipeline next step (one line)

lineups.filter(pl.col("mins") > 10).sort("netrtg", descending=True).head()

ncaa_mbb_on_off​

ncaa_mbb_on_off(players: 'Union[str, Sequence[str]]', lineups: 'pl.DataFrame', *, included: 'Union[str, Sequence[str], None]' = None, excluded: 'Union[str, Sequence[str], None]' = None, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Team stats for every on/off combination of the given players.

Port of bigballR on_off_generator (all_functions.R:2555-2749, include_transition=F path). For k players, all 2^k on/off assignments are enumerated in R's expand.grid order (first player varies fastest; first row all-On, last all-Off); each combination sums the lineups whose membership matches exactly, re-derives the full ratio block (including e_poss), rounds to 3 decimals, and zeroes NA/Inf. Combinations matching zero lineups produce an all-zero row.

Faithful R quirk (kept, flagged): when included/excluded is passed, the base lineup set comes from the membership filter ONLY — the inferred-team filter is skipped (all_functions.R:2584-2589), so an included player on another team would leak that team's lineups in.

Parameters

ParameterTypeDefaultDescription
playersUnion[str, Sequence[str]]Player name(s) to split on (the Status axis).
lineupsDataFrameLineups frame from ncaa_mbb_lineups.
includedUnion[str, Sequence[str], None]NoneOptional membership filter forwarded to ncaa_mbb_player_lineups (replaces the team filter).
excludedUnion[str, Sequence[str], None]NoneOptional membership filter forwarded to ncaa_mbb_player_lineups (replaces the team filter).
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with 2^k rows — status (e.g. "A.PLAYER On | B.PLAYER Off") + the 69 stat columns (ON_OFF_COLUMNS).

Example

from sportsdataverse.mbb.mbb_ncaa_lineups import ncaa_mbb_on_off
split = ncaa_mbb_on_off("KEATON.WAGLER", lineups)
print(split.shape)

# Two-player interaction

duo = ncaa_mbb_on_off(["A.PLAYER", "B.PLAYER"], lineups)

# Pipeline next step (one line)

split.select("status", "netrtg")

ncaa_mbb_player_combos​

ncaa_mbb_player_combos(lineups: 'pl.DataFrame', *, n: 'int' = 2, min_mins: 'float' = 0, included: 'Union[str, Sequence[str], None]' = None, excluded: 'Union[str, Sequence[str], None]' = None, include_transition: 'bool' = False, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Team stats for every n-player combination on the court together.

Port of bigballR get_player_combos + team_comb (get_player_combos.R:19-38 + :42-191). Combos are enumerated per team over the byte-sorted unique player pool (lexicographic gtools::combinations order), filtered to combos whose lineups total strictly more than min_mins minutes, then each combo's lineup rows are summed and the ratio block re-derived. Rounding is 2 decimals here (R rounds 3 in get_lineups / on_off_generator) and e_poss is the SUM of the per-lineup estimates, not a recompute — both faithful.

R's include_transition switch is dead code (an exact-match guard at get_player_combos.R:28-30 always forces it back to FALSE); the port implements the suffix-match intent but keeps the False default, which is the only R-reachable behavior.

Parameters

ParameterTypeDefaultDescription
lineupsDataFrameLineups frame from ncaa_mbb_lineups.
nint2Combination size, 1-5.
min_minsfloat0Keep combos with total on-court minutes strictly greater than this (summed over the rounded per-lineup mins).
includedUnion[str, Sequence[str], None]NonePlayer name(s) that must be on the court in every lineup considered.
excludedUnion[str, Sequence[str], None]NonePlayer name(s) that must be off the court in every lineup considered.
include_transitionboolFalseRe-derive the trans/half ratio surface (requires a transition lineups frame; forced False otherwise).
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with one row per combo: team, p1..pn + the stat surface of the input frame. Teams appear in byte-sorted order.

Example

from sportsdataverse.mbb.mbb_ncaa_lineups import ncaa_mbb_player_combos
duos = ncaa_mbb_player_combos(lineups, n=2, min_mins=5)
print(duos.shape)

# Anchored on one player

trios = ncaa_mbb_player_combos(lineups, n=3, included="KEATON.WAGLER")

# Pipeline next step (one line)

duos.sort("netrtg", descending=True).head()

ncaa_mbb_player_lineups​

ncaa_mbb_player_lineups(lineups: 'pl.DataFrame', *, included: 'Union[str, Sequence[str], None]' = None, excluded: 'Union[str, Sequence[str], None]' = None, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Filter a lineups frame by on-court player membership.

Port of bigballR get_player_lineups (all_functions.R:2761-2792): keep rows where every included player is on the court (in p1..p5) and no excluded player is. With both filters None the input is returned unchanged (R's Included = NA, Excluded = NA passthrough). Membership is tested by name against p1..p5 (R tests the positional first five columns); row order is preserved.

Parameters

ParameterTypeDefaultDescription
lineupsDataFrameLineups frame from ncaa_mbb_lineups (any frame with p1..p5 works).
includedUnion[str, Sequence[str], None]NonePlayer name(s) that must ALL be on the court.
excludedUnion[str, Sequence[str], None]NonePlayer name(s) that must NONE be on the court.
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

Row-subset of lineups; schema unchanged.

Example

from sportsdataverse.mbb.mbb_ncaa_lineups import ncaa_mbb_player_lineups
on = ncaa_mbb_player_lineups(lineups, included="KEATON.WAGLER")
print(on.shape)

# Included + excluded combination

df = ncaa_mbb_player_lineups(lineups, included=["A.PLAYER"], excluded=["B.PLAYER"])

# Pipeline next step (one line)

on.select(pl.col("mins").sum())

ncaa_mbb_player_stats​

ncaa_mbb_player_stats(pbp: 'pl.DataFrame', *, multi_games: 'bool' = False, simple: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into per-player box stats.

Port of bigballR get_player_stats (all_functions.R:2810-3177) + its get_mins helper (:3240-3263). Counting stats are summarised per (game, team, player), assists counted from player_2, minutes and offensive possessions derived from the ten on-court columns, and rates (FG%, TS%, eFG%, rim/mid splits, ...) computed from the counters and rounded to 3 decimals with R's round semantics. With multi_games=True the per-game rows are summed per (player, team), every rate is recomputed from the summed counters (never averaged), and GP/GS are appended.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract. May span multiple games.
multi_gamesboolFalseWhen True, aggregate across games per (player, team) — the season-stat surface. When False (default, R parity), treat each game separately and keep the game id columns.
simpleboolFalseWhen True, return the reduced 33-column (multi) / 35-column (per-game) surface without the transition / assisted / putback / block-location splits.
fix_tip_inboolTrueWhen True (default), rim and putback stats count the scrape engine's real "Tip In" vocabulary. When False, reproduce R's literal "Tip-In" test (all_functions.R:2827) — tip-ins silently excluded — for oracle parity.
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame): one row per player+team (+game when multi_games=False). Columns follow PLAYER_STATS_COLUMNS / PLAYER_STATS_SIMPLE_COLUMNS / PLAYER_GAME_STATS_COLUMNS / PLAYER_GAME_STATS_SIMPLE_COLUMNS. Rows sorted by the group keys (byte order, matching dplyr's C-locale group order). Empty input yields an empty frame with the documented schema.

Example

from sportsdataverse.mbb.mbb_ncaa_stats_agg import ncaa_mbb_player_stats
season = ncaa_mbb_player_stats(pbp, multi_games=True)
print(season.shape)

# Reduced surface, pandas out

df_pd = ncaa_mbb_player_stats(pbp, multi_games=True, simple=True, return_as_pandas=True)

# Pipeline next step (one line)

season.filter(pl.col("mins") > 50).sort("pts", descending=True).head()

ncaa_mbb_possessions​

ncaa_mbb_possessions(pbp: 'pl.DataFrame', *, simple: 'bool' = False, fix_cross_game_leak: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into one row per possession.

Port of bigballR get_possessions (all_functions.R:3686-3745). Groups by the possession keys stamped upstream by the scrape engine (poss_num, poss_team, the ten on-court lineup columns, plus game identity), drops possessions with any missing on-court player, and — in the full variant — sorts each row's home/away lineup alphabetically so a given lineup always occupies the same columns.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (parse_ncaa_bb_game_pbp output). May span multiple games; rows must be in scrape order.
simpleboolFalseWhen True, return only the 17-column possession/points frame (all_functions.R:3687-3694) with lineups in on-court order. When False (default), return the full 28-column frame with per-possession context columns and alpha-sorted lineups.
fix_cross_game_leakboolTrueWhen True (default, and the CORRECT behavior), window the start_event_type lag with .over("game_id") so a game's first possession has a null start event instead of inheriting the PREVIOUS game's last event. When False, reproduce R's ungrouped dplyr::lag (all_functions.R:3698) and its cross-game leak. Parity tests pass False. Ignored when simple=True (that variant emits no start_event_type).
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with one row per possession — 28 columns per POSSESSION_SEG_SCHEMA (full) or 17 per POSSESSIONS_SIMPLE_SCHEMA (simple). Empty input yields an empty frame carrying the documented schema.

Example

from sportsdataverse.mbb.mbb_ncaa_possession_seg import ncaa_mbb_possessions
poss = ncaa_mbb_possessions(pbp)
print(poss.shape)

# Simple points-per-possession variant

poss_pd = ncaa_mbb_possessions(pbp, simple=True, return_as_pandas=True)

# Pipeline next step (one line)

poss.group_by("poss_team").agg(pl.col("pts").mean())

ncaa_mbb_shot_locations​

ncaa_mbb_shot_locations(game_ids: "'Sequence[object]'", *, fetcher: 'Optional[_SupportsFetchGameBox]' = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, Any]'"

Scrape MBB shot locations for one or more games (bigballR

get_shot_locations, get_shot_locations.R:3-89).

Fetches each game's stats.ncaa.org/contests/{id}/box_score page and parses the embedded shot-chart JS through parse_ncaa_bb_shots. NA ids are dropped up front (R :5); per-game "shots found" messages go to the module logger (R message, :69-70).

Parameters

ParameterTypeDefaultDescription
game_idsSequence[object]NCAA contest ids; None/NaN entries are dropped.
fetcherOptional[_SupportsFetchGameBox]NoneOptional injected fetcher exposing fetch_game_box (for tests/offline use). Defaults to a fresh NcaaFetcher.with_browser() context per call.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

All games' shots row-bound (zero-row SHOTS_SCHEMA frame when no ids survive or no charts are found).

Example

from sportsdataverse.mbb.mbb_ncaa_shots import ncaa_mbb_shot_locations
df = ncaa_mbb_shot_locations(["6470186", "6479639"])
print(df.shape)

# Offline with an injected fetcher

df = ncaa_mbb_shot_locations(["6470186"], fetcher=my_fetcher)

# Pipeline next step (one line)

df.group_by("team").agg(pl.col("shot_dist").mean()).head()

ncaa_mbb_team_ids​

ncaa_mbb_team_ids(*, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Men's-basketball (team, season) -> stats.ncaa.org id crosswalk.

Port of bigballR's bundled teamids data asset (one row per team per season, 2009-10 through 2025-26).

Parameters

ParameterTypeDefaultDescription
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

DataFrame with columns team (str), conference (str), id (Int64 -- the season-specific stats.ncaa.org team id) and season (str, "YYYY-YY").

Example

from sportsdataverse.mbb import ncaa_mbb_team_ids
df = ncaa_mbb_team_ids()
print(df.shape)

# Pipeline next step (one line)

df.filter(pl.col("season") == "2025-26").head()

ncaa_mbb_team_stats​

ncaa_mbb_team_stats(pbp: 'pl.DataFrame', *, include_transition: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into per-team game stats.

Port of bigballR get_team_stats (all_functions.R:2530-2538): the ten on-court columns are blanked so every row shares one "lineup", then get_lineups (ncaa_mbb_lineups) runs per game and the lineup key columns are dropped — yielding two rows (one per team) per game.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract. May span multiple games.
include_transitionboolFalseWhen True, append the trans/half split surface plus o_trans_pct/d_trans_pct.
fix_tip_inboolTrueWhen True (default), rim stats count the scrape engine's real "Tip In" vocabulary; False reproduces R's literal "Tip-In" bug for oracle parity.
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with one row per team per game — TEAM_STATS_COLUMNS (73) or TEAM_STATS_TRANSITION_COLUMNS with include_transition=True. Games ordered by the Utf8 game_id byte sort (R's do() sorts a numeric ID — identical for equal-width ids), teams within a game byte-sorted. Empty input yields an empty frame with the documented schema.

Example

from sportsdataverse.mbb.mbb_ncaa_stats_agg import ncaa_mbb_team_stats
teams = ncaa_mbb_team_stats(pbp)
print(teams.shape)

# Transition splits, pandas out

df_pd = ncaa_mbb_team_stats(pbp, include_transition=True, return_as_pandas=True)

# Pipeline next step (one line)

teams.sort("netrtg", descending=True).head()