Skip to main content

WBB — additional Python functions — stats.ncaa.org: enrich_stats–phase1_shot

enrich_stats​

enrich_stats(lineup: 'LineupEvent', event_parser: 'PossessionEvent', stats: 'LineupEventStats', player_filter_coder: 'Optional[PlayerFilterCoder]' = None, player_index: 'int' = -1) -> 'LineupEventStats'

Fold a lineup's raw events into a counting-stat tree (``protected def

enrich_stats, LineupUtils.scala:115-162``). Reuses the Task 5a.3 concurrent-clump batching (~sportsdataverse.mbb.mbb_ncaa_possessions .lineup_as_raw_clumps + ~sportsdataverse.mbb.mbb_ncaa_possessions .concurrent_event_handler) rather than duplicating it -- both were already public/exported from Task 5a.3.

stats is deep-copied once up front (see the module docstring's "Scala idiom decisions"), so this function never mutates the caller's stats argument -- safe to call repeatedly against the same starting literal (e.g. a shared "empty stats" fixture).

Parameters

ParameterTypeDefaultDescription
lineupLineupEventThe lineup whose raw_game_events to fold over.
event_parserPossessionEventSelects which side (team/opponent) is "attacking".
statsLineupEventStatsThe starting stat tree (not mutated -- see above).
player_filter_coderOptional[PlayerFilterCoder]NoneOptional name -> (is_this_player, code) predicate/coder, for per-player scoping (Task 5c.4).
player_indexint-1Lineup-slot index for ~sportsdataverse.mbb .mbb_ncaa_models.PlayerShotInfo tuples (Task 5c.4; -1 for team-level calls, the only value exercised before then).

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every matching event folded in.

ensure_ev_uniqueness​

ensure_ev_uniqueness(clump: 'ConcurrentClump') -> 'ConcurrentClump'

Nudge each event's min by a tiny per-index delta so truly

concurrent (identical-min) events within a clump don't collapse under == (ensure_ev_uniqueness, LineupUtils.scala:105-111).

Parameters

ParameterTypeDefaultDescription
clumpConcurrentClumpThe clump whose events to nudge.

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_possessions.ConcurrentClump with each event's min incremented by 1e-6 * index.

extract_player_from_ev​

extract_player_from_ev(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', tidy_ctx: 'TidyPlayerContext') -> 'Optional[PlayerCodeId]'

Resolve the player named in pbp_event to a

~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (ShotEnrichmentUtils.extract_player_from_ev, PlayByPlayUtils.scala:613-635).

For a shot by the team under analysis (shot.is_off) the name is tidied against the box score before coding (so a mis-spelled play-by-play name resolves to the roster identity); for an opponent shot it is coded verbatim with no team context.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched (only is_off is read).
pbp_eventMiscGameEventThe play-by-play event naming the player.
tidy_ctxTidyPlayerContextThe name-resolution context for this game.

Returns

The resolved PlayerCodeId, or None if the event string names no player (~sportsdataverse.mbb.mbb_ncaa_events.parse_any_play found nothing).

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import extract_player_from_ev
pc = extract_player_from_ev(shot, pbp_event, tidy_ctx)

field_keys​

field_keys(field: 'str') -> 'dict[str, str]'

Off/def stat-key names for a field (fieldKeys, ts:77-79).

Parameters

ParameterTypeDefaultDescription
fieldstrA stat field ("efg" / "3p" / "2pmid" / "2prim").

Returns

{"off": f"off_{field}", "def": f"def_{field}"}.

Example

from sportsdataverse.mbb.mbb_ncaa_strength import field_keys

keys = field_keys("3p")
print(keys["off"], keys["def"]) # off_3p def_3p

filter_matching_own​

filter_matching_own(tags: 'list[Tag]', regex: 'str') -> 'list[Tag]'

JSoup :matchesOwn(regex) applied to an already-computed candidate

list, rather than a fresh root.select(selector) call (Task 5e.2 addition; see the module docstring's note on composing this with attr_regex_filter).

Same own-text-only semantics as select_matching_own -- JSoup's Element.ownText() walks only the element's direct TextNode children, not text nested inside child elements.

Parameters

ParameterTypeDefaultDescription
tagslist[Tag]Candidate tags to filter (typically the result of an earlier .select()/attr_regex_filter call).
regexstrThe pattern each candidate's own (whitespace-collapsed) text must re.search-match.

Returns

The subset of tags whose own text contains a regex match, in the input list's order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import attr_regex_filter, filter_matching_own, parse_html
soup = parse_html('<td style="font-size:36px">92</td><td style="color:red">x</td>')
candidates = attr_regex_filter(soup.find_all("td"), "style", r"font-size:36px")
filter_matching_own(candidates, r"[0-9]+") # [<td style="font-size:36px">92</td>]

find_lineup​

find_lineup(shot: 'ShotEvent', curr_pbp: 'Optional[MiscGameEvent]', curr_lineups: 'list[LineupEvent]', lineup_it: "'PeekableIterator[LineupEvent]'") -> 'tuple[Optional[LineupEvent], list[LineupEvent]]'

Find the lineup (stint) event on the floor for shot

(ShotEnrichmentUtils.find_lineup, PlayByPlayUtils.scala:352-517).

A recursive state machine over three lists: curr_lineup (the current candidate), fallback_lineups (time-matching lineups whose raw events did not contain curr_pbp -- kept as fallbacks), and stashed_lineups (lineups pulled from the iterator but not yet stepped into, available for future shots). The branch cases (labelled 2.1-2.4 in the Scala):

  • 2.1 -- no time-matching lineup left: return the fallbacks.
  • 2.2 -- the next lineup starts after the shot: no match, stash it.
  • 2.3 -- strictly inside a lineup with no prior fallbacks: take it.
  • 2.4 -- shot is exactly at a lineup boundary (or we are already choosing among multiple fallbacks): take this lineup iff its raw game events contain curr_pbp's event string (curr_pbp is None takes it unconditionally); otherwise keep it as a fallback and recurse.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot to place (only min / is_off are read).
curr_pbpOptional[MiscGameEvent]The already-matched play-by-play event for this shot, used to disambiguate boundary lineups; None disables that check.
curr_lineupslist[LineupEvent]Lineups pulled from the iterator on a previous call and still available (the current one first).
lineup_itPeekableIterator[LineupEvent]The shared lineup iterator (consumed in place).

Returns

(matched_lineup_or_None, lineups_to_retry_next_time) -- the second element always includes the matched lineup (so out-of-order shots sharing it still resolve) plus any leftover stash.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import (
PeekableIterator,
find_lineup,
)
matched, retry = find_lineup(shot, None, [lineup], PeekableIterator([]))

find_missing_subs​

find_missing_subs(clump: 'BadLineupClump', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'tuple[list[LineupEvent], BadLineupClump]'

Trims a clump whose lineups carry TOO MANY players by identifying the

"ghost" player(s) a missing sub-out left behind (LineupErrorAnalysisUtils.find_missing_subs, :406-514).

Fires only when the clump's first event has >= 6 on-floor players (:415-416: candidates.size < 6 is a no-op). expected_size_diff (:419) is first_event_player_count - 5 -- the number of ghosts the trim should end up removing.

Phase 1 -- shrink the candidate pool (:437-478). Starting from the first event's players, walk the clump chronologically. At each event a candidate is confirmed present (and dropped from the pool) if it subs out (ev.players_out, skipped for the first event -- :445, literal port of clump.evs.headOption.contains(ev) as value equality ev == clump.evs[0]; for a well-formed clump of distinct events this is exactly index == 0) or is named in one of the event's team-side raw plays (parse_any_play -> ~sportsdataverse.mbb.mbb_ncaa_names .tidy_player -> ~sportsdataverse.mbb.mbb_ncaa_stints .build_player_code, :448-456; unlike validate_lineup this does NOT skip the literal "team" token -- ported verbatim). matching_index is the FIRST event index at which the pool size first equals expected_size_diff (:475); once set it freezes -- all later events are skipped in phase 1 (:439-441).

Accept gate (:479-480): the final pool must be non-empty and no larger than expected_size_diff. If matching_index never fired (the pool jumped past expected_size_diff in a single step, or never shrank to it), the gate still accepts iff the residual pool is a non-empty subset of size <= expected_size_diff -- in which case phase 3 routes every event through the "before match" branch (index > None is always false). On failure the original clump is returned unchanged.

Phase 3 -- rebuild the events (:482-503, a scanLeft ported as a manual accumulate loop that drops the seed). For events at/before matching_index the ghost pool is simply removed from players (filterNot). For events strictly after matching_index (index > matching_index -- the matched event itself is "before") players is rebuilt from the previous tidied event via ~sportsdataverse.mbb.mbb_ncaa_stints.build_new_player_list (the scanLeft threads the previously-emitted event; its seed is None, but the first event can never be an "after match" event, so the getOrElse(ev) fallback is only ever a formality -- ported faithfully all the same). The rebuilt events are partitioned by validate_lineup.

Parameters

ParameterTypeDefaultDescription
clumpBadLineupClumpThe bad-lineup clump to attempt to repair.
box_lineupLineupEventThe team's box-score lineup event (roster + name context).
valid_player_codesset[str]Every player code on the box score / roster.

Returns

(fixed_lineups, still_to_fix) -- the now-valid rebuilt events and a clump of the still-invalid ones (carrying the input's next_good); or ([], clump) on a no-op / rejected fix.

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import (
find_missing_subs,
)
fixed, still = find_missing_subs(clump, box_lineup, valid_codes)

find_pbp_clump​

find_pbp_clump(shot_time: 'float', pbp_it: "'PeekableIterator[PlayByPlayEvent]'", curr_pbp_clump: 'list[MiscGameEvent]', maybe_next_pbp_event: 'Optional[MiscGameEvent]') -> 'tuple[list[MiscGameEvent], Optional[MiscGameEvent]]'

Gather every play-by-play shot/assist event sharing shot_time

(ShotEnrichmentUtils.find_pbp_clump, PlayByPlayUtils.scala:556-608).

If curr_pbp_clump (carried over from the previous shot) already holds events at shot_time they are returned as-is; otherwise the iterator is walked forward, discarding earlier events, accumulating the equal-time ones, and stopping (returning it as maybe_next_pbp_event) at the first later event.

Parameters

ParameterTypeDefaultDescription
shot_timefloatThe shot's game-clock minute to gather events for.
pbp_itPeekableIterator[PlayByPlayEvent]The shared play-by-play iterator (consumed in place).
curr_pbp_clumplist[MiscGameEvent]Events left over from the previous shot's clump.
maybe_next_pbp_eventOptional[MiscGameEvent]The look-ahead event stashed by the previous call, if any.

Returns

(clump, maybe_next) -- the equal-time events, plus the first strictly-later event (or None at end of stream).

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import (
PeekableIterator,
find_pbp_clump,
)
clump, nxt = find_pbp_clump(5.0, PeekableIterator([]), [], None)
# ([], None)

fix_combos​

fix_combos(first: 'str', last: 'str', code_start: 'Optional[str]' = None) -> 'list[tuple[str, Optional[str]]]'

Pair each of combos' three name variants with a shared

player-code override (DataQualityIssues.fix_combos, DataQualityIssues.scala:340-346).

Parameters

ParameterTypeDefaultDescription
firststrThe player's first name.
laststrThe player's last name.
code_startOptional[str]NoneThe forced player-code prefix for every variant, or None to leave the default build_player_code truncation behavior in place.

Returns

Three (name_variant, code_start) pairs.

fix_possible_score_swap_bug​

fix_possible_score_swap_bug(lineup: 'list[LineupEvent]', box_lineup: 'LineupEvent') -> 'list[LineupEvent]'

Undo a rare NCAA data bug where the scores get transposed

(fix_possible_score_swap_bug, LineupUtils.scala:51-90).

If the last lineup's ending score is the exact transpose of the box score's ending score, every lineup's score_info is un-transposed and pts/plus_minus are swapped/negated between team_stats and opponent_stats -- nothing else in the stat trees changes.

Parameters

ParameterTypeDefaultDescription
lineuplist[LineupEvent]The lineups to (maybe) fix, in chronological order.
box_lineupLineupEventThe trusted box-score lineup to compare the final score against.

Returns

lineup unchanged if the scores aren't transposed (or lineup is empty); otherwise a new list with every entry's score/pts/ plus_minus corrected.

get_ascending_time​

get_ascending_time(event: 'ShotEvent', period: 'int', is_women_game: 'bool') -> 'float'

Converts the descending in-period clock time to an ascending

game-elapsed time (ShotEventParser.get_ascending_time, :531-537).

Parameters

ParameterTypeDefaultDescription
eventShotEventThe shot event (only ~sportsdataverse.mbb .mbb_ncaa_models.ShotEvent.min, the raw descending clock minute, is read).
periodintThe 1-indexed period the shot was taken in.
is_women_gameboolWhether to use women's-quarters (10min) or men's- halves (20min, then 5min OTs) period lengths.

Returns

The ascending game-elapsed time, in minutes.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import get_ascending_time
get_ascending_time(shot_with_min_4, period=1, is_women_game=False) # 16.0

get_box_lineup​

get_box_lineup(filename: 'str', in_html: 'str', team_id: 'TeamId', format_version: 'int', external_roster: 'tuple[list[str], list[RosterEntry]]' = ([], []), neutral_game_dates: 'AbstractSet[str]' = frozenset(), home_team: 'Optional[str]' = None, away_team: 'Optional[str]' = None) -> 'Union[LineupEvent, list[ParseError]]'

Gets the boxscore lineup from the HTML page (``BoxscoreParser

.get_box_lineup, :122-222``).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name -- used both for error reporting and to extract the period via parse_period_from_filename (e.g. "test_p2.html").
in_htmlstrThe raw box-score-page HTML.
team_idTeamIdThe team this box score is being parsed for.
format_versionint0 for the legacy layout, 1 for the 2018+ layout (see the module docstring's selector-translation notes).
external_rostertuple[list[str], list[RosterEntry]]([], [])(other_players, roster_players) -- either just names, or a full roster, to validate/fuzzy-correct box names against (see inject_validated_players). Also seeds ~sportsdataverse.mbb.mbb_ncaa_models.LineupEvent.players_out on the interim lineup (roster_players, each's code replaced by its jersey number).
neutral_game_datesAbstractSet[str]frozenset()Date strings (the first whitespace-separated token of the raw date-cell text) known to be neutral-site games -- overrides the default home/away inference.
home_teamOptional[str]NoneThe game's home team when the caller already knows it, forwarded to team-name resolution so a box page that names only one side (a non-D-I opponent has no header) still resolves.
away_teamOptional[str]NoneThe game's away team, same purpose. Both are required together or neither is used.

Returns

A ~sportsdataverse.mbb.mbb_ncaa_models.LineupEvent whose players is the validated box-score lineup (natural HTML order -- see the module docstring's "not sorted" note), or a list[ParseError] if any parsing step failed.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import get_box_lineup
from sportsdataverse.mbb.mbb_ncaa_models import TeamId

with open("tests/fixtures/ncaa/test_lineup.html", encoding="utf-8") as f:
html = f.read()
result = get_box_lineup("test_p1.html", html, TeamId("TeamA"), format_version=0)

get_config​

get_config() -> 'NcaaFetchConfig'

Return the live NcaaFetchConfig singleton.

Returns

The live singleton (cache_dir, proxy_url, the ProxyBonanza settings, timeout, impersonate, max_retries and the rotation / Terms-gate backoffs, transport).

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import get_config
cfg = get_config()
print(cfg.cache_dir, cfg.timeout)

get_game_weight​

get_game_weight(opp: 'OpponentGame', field: 'str', side: 'str') -> 'float'

Weight for one game/field/side (getGameWeight, ts:119-140).

The field-specific shot volume (FGA for efg, 3PA for 3p, 2pmid_attempts / 2prim_attempts for the mid/rim fields); when that is 0 (no shots of that type), falls back to off_poss / def_poss so the game still carries weight.

Parameters

ParameterTypeDefaultDescription
oppOpponentGameOne opponent game dict.
fieldstrA stat field.
sidestr"off" or "def".

Returns

The (non-negative) game weight.

Example

from sportsdataverse.mbb.mbb_ncaa_strength import get_game_weight

game = {"off_3p_attempts": 0, "off_poss": 70}
print(get_game_weight(game, "3p", "off")) # 70.0 (poss fallback)

get_neutral_games​

get_neutral_games(filename: 'str', in_html: 'str', format_version: 'int') -> 'Union[tuple[TeamId, set[str]], list[ParseError]]'

Extracts the set of neutral/away-marked game dates from a saved NCAA

team-schedule page (TeamScheduleParser.get_neutral_games, TeamScheduleParser.scala:63-94).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw team-schedule-page HTML.
format_versionint0 for the legacy fieldset/legend layout, 1 for the 2018+ div.card-header/div.card-body layout.

Returns

(team, neutral_game_dates) -- the team parsed from the page's image alt attribute, and every "MM/DD/YYYY" date string found on an "@Opponent"-marked row -- or a single-element list[ParseError] if the HTML fails to parse, or the team name can't be located.

Example

from sportsdataverse.mbb.mbb_ncaa_team_parsers import get_neutral_games

with open("tests/fixtures/ncaa/test_schedule.html", encoding="utf-8") as f:
html = f.read()
result = get_neutral_games("test_schedule.html", html, format_version=0)
if isinstance(result, list):
raise RuntimeError(result) # list[ParseError]
team, neutral_dates = result

get_per_game_raw​

get_per_game_raw(opp: 'OpponentGame', field: 'str', side: 'str') -> 'Optional[float]'

Per-game raw shooting rate from one opponent row (getPerGameRaw, ts:82-116).

efg is (2pmid_made + 2prim_made + 1.5 * 3p_made) / (2pmid_att + 2prim_att + 3p_att); 3p / 2pmid / 2prim are made / attempts. Every counter read is nullish (missing -> 0); the sole guard is on total attempts.

Parameters

ParameterTypeDefaultDescription
oppOpponentGameOne opponent game dict.
fieldstrA stat field; an unknown field returns None.
sidestr"off" or "def" (selects the off_/def_ prefix).

Returns

The rate as a float, or None when the relevant attempts total is <= 0 (game skipped by the weighted means -- not a 0-rate).

Example

from sportsdataverse.mbb.mbb_ncaa_strength import get_per_game_raw

game = {"off_3p_made": 4, "off_3p_attempts": 10}
print(get_per_game_raw(game, "3p", "off")) # 0.4

get_sorted_pbp_events​

get_sorted_pbp_events(filename: 'str', in_html: 'str', box_lineup: 'LineupEvent', format_version: 'int') -> 'Union[list[PlayByPlayEvent], list[ParseError]]'

Handy util to return the play-by-play events in chronological order,

used in a few other places (PlayByPlayParser.get_sorted_pbp_events, :221-239).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw play-by-play-page HTML.
box_lineupLineupEventThe team's box-score lineup (supplies team/year).
format_versionint0 for the legacy layout, 1 for the 2018+ layout.

Returns

The play-by-play events in chronological (earliest-to-latest) order, or a list[ParseError] on failure. enrich=True is used internally to get the correct ascending timestamps, and its reversal is undone here (.reverse) to restore chronological order.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import get_box_lineup
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
from sportsdataverse.mbb.mbb_ncaa_pbp_parser import get_sorted_pbp_events

with open("tests/fixtures/ncaa/test_lineup.html", encoding="utf-8") as f:
box_html = f.read()
box_lineup = get_box_lineup("test_p1.html", box_html, TeamId("TeamA"), format_version=0)

with open("tests/fixtures/ncaa/test_play_by_play.html", encoding="utf-8") as f:
pbp_html = f.read()
events = get_sorted_pbp_events("test.html", pbp_html, box_lineup, format_version=0)

get_team_raw_from_per_game​

get_team_raw_from_per_game(team: 'TeamDetail', field: 'str') -> 'SideValues'

A team's field rate as the weighted mean of its per-game raws (getTeamRawFromPerGame, ts:224-250).

Same accumulation as compute_league_averages_from_per_game but scoped to one team's games; empty -> 0.

Parameters

ParameterTypeDefaultDescription
teamTeamDetailA team_details team dict.
fieldstrA stat field.

Returns

{"off": float, "def": float}.

Example

from sportsdataverse.mbb.mbb_ncaa_strength import get_team_raw_from_per_game

team = {"opponents": [{"off_3p_made": 4, "off_3p_attempts": 10}]}
print(get_team_raw_from_per_game(team, "3p")["off"]) # 0.4

get_team_triples​

get_team_triples(filename: 'str', in_html: 'str', old_format: 'bool' = False) -> 'Union[list[tuple[TeamId, str, ConferenceId]], list[ParseError]]'

Extracts (team, NCAA id, conference) triples from a saved NCAA

team-list/attendance page (TeamIdParser.get_team_triples, TeamIdParser.scala:69-91).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw team-list-page HTML.
old_formatboolFalseTrue for pages where the team name and conference are in separate <td>s; False (default) for pages where the conference is embedded in the team-name cell as "Team (Conf)".

Returns

One (TeamId, ncaa_id, ConferenceId) triple per row that has both a resolvable id and name/conference (rows missing either are silently skipped, matching the Scala's case _ => Nil), or a single-element list[ParseError] if the HTML itself fails to parse.

Example

from sportsdataverse.mbb.mbb_ncaa_team_parsers import get_team_triples

with open("tests/fixtures/ncaa/test_attendance_list.html", encoding="utf-8") as f:
html = f.read()
result = get_team_triples("test_attendance_list.html", html, old_format=True)

get_unified_ncaa_id​

get_unified_ncaa_id(filename: 'str', in_html: 'str') -> 'Union[Optional[str], list[ParseError]]'

Gets a player's lowest cross-season NCAA id from a saved player page

(RosterParser.get_unified_ncaa_id, RosterParser.scala:136-152).

Always uses the v1 selector table -- this bonus lookup only exists on 2018+-era pages.

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw player-page HTML.

Returns

The numerically-lowest NCAA id found, None if the page has no tr[id^=player_season_] rows, or a single-element list[ParseError] if the HTML couldn't be parsed at all.

Example

from sportsdataverse.mbb.mbb_ncaa_roster_parser import get_unified_ncaa_id
get_unified_ncaa_id("player.html", player_page_html)

handle_common_sub_bug​

handle_common_sub_bug(clump: 'BadLineupClump', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'tuple[list[LineupEvent], BadLineupClump]'

Fixes the "2-in-1-out then a compensating 1-out" substitution bug

(LineupErrorAnalysisUtils.handle_common_sub_bug, :269-298).

Handles a single-event bad clump whose next known-good lineup carries a lone sub-out that the clump's event should have applied but didn't (e.g. IN: X, Y, Z; OUT: A, B in the bad event, then OUT: C in the good one). Fires only when all three guard conditions hold (:275-278):

  • the bad event has more players subbing IN than OUT (len(players_in) > len(players_out)),
  • the good event has no sub-ins (len(good.players_in) == 0 -- otherwise there's no way to tell which of its sub-outs to borrow), and
  • the good event has at least one sub-out (len(good.players_out) > 0).

The fix (:279-283) removes the good event's sub-outs from the bad event's on-floor players (value-equality not in -- the Scala's filterNot(good.players_out.toSet)) and appends them to the bad event's players_out (order-preserving distinct-- the Scala's(bad.players_out ++ good.players_out).distinct). The fix is **accepted only if** the result then passes validate_lineup (:284-294): on success the fixed event is returned as the sole fixedlineup and the still-to-fix clump is emptied; on failure the *fixed* event (not the original) is returned as the still-to-fix clump, keeping the samenext_good` so a later pass can try again.

Parameters

ParameterTypeDefaultDescription
clumpBadLineupClumpThe bad-lineup clump to attempt to repair.
box_lineupLineupEventThe team's box-score lineup event (roster + name context).
valid_player_codesset[str]Every player code on the box score / roster.

Returns

(fixed_lineups, still_to_fix) -- fixed_lineups is [fixed] on an accepted fix else []; still_to_fix is an empty clump on accept, the (unchanged) input clump on a guard miss, or the single-event fixed-but-still-invalid clump on a rejected fix.

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import (
handle_common_sub_bug,
)
fixed, still = handle_common_sub_bug(clump, box_lineup, valid_codes)

inject_starting_lineup_into_box​

inject_starting_lineup_into_box(sorted_pbp_events: 'list[PlayByPlayEvent]', box_lineup: 'LineupEvent', external_roster: 'tuple[list[str], list[RosterEntry]]', format_version: 'int') -> 'LineupEvent'

Infer the starting five and reorder the box-score roster so they lead

(PlayByPlayUtils.inject_starting_lineup_into_box, PlayByPlayUtils.scala:684-845).

The v1 (2018+) NCAA box score dropped the ordered list of starters, so we reconstruct it from the play-by-play sub sequencing. A player is a starter if, walking the events forward, they are seen before their first sub-in -- either subbed out (before ever being subbed in), or named in a team-side play that isn't concurrent with a sub. Anyone subbed in before ever being seen is excluded. The reconstruction stops once five starters are found.

Parameters

ParameterTypeDefaultDescription
sorted_pbp_eventslist[PlayByPlayEvent]The full play-by-play event stream, ascending time.
box_lineupLineupEventThe box-score lineup event (its players is the full roster to reorder).
external_rostertuple[list[str], list[RosterEntry]]Unused here -- carried for signature parity with the Scala (its pipeline caller passes it). See the module note.
format_versionintUnused here -- carried for signature parity. See the module note.

Returns

A copy of box_lineup with players reordered so the inferred starters lead. If fewer than five starters could be inferred (a "40-trillion" player who was never subbed nor mentioned), the roster is ordered starters -> possible-starters -> definitely-not-starters as the best available guess.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import inject_starting_lineup_into_box
fixed = inject_starting_lineup_into_box(pbp_events, box_lineup, ([], []), 1)

inject_validated_players​

inject_validated_players(ordered_lineup_from_box: 'list[str]', box_minus_players: 'LineupEvent', external_roster: 'tuple[list[str], list[RosterEntry]]') -> 'list[str]'

Validates box players against the roster (if available) and any

other available box scores (BoxscoreParser.inject_validated_players, :233-279).

See the module docstring's "un-threaded tidy_ctx" note -- every fuzzy-resolution call inside the loop uses the SAME original context, never the updated one a call returns (ported verbatim, including this apparent Scala oversight).

Parameters

ParameterTypeDefaultDescription
ordered_lineup_from_boxlist[str]The raw player-name strings scraped straight off the box-score page (already v0-normalized if the source was v1, by get_box_lineup's caller).
box_minus_playersLineupEventThe in-progress ~sportsdataverse.mbb.mbb_ncaa_models.LineupEvent (used only for its team field, both to scope the fuzzy-match context and to key ~sportsdataverse.mbb.mbb_ncaa_data_quality.players_missing_from_boxscore).
external_rostertuple[list[str], list[RosterEntry]](other_players, roster_players) -- extra known player names, and a full team roster (if available) to validate against / fuzzy-correct box names onto.

Returns

ordered_lineup_from_box with any name not found in roster_players fuzzy-corrected onto the closest roster name (if a roster was supplied at all), followed by any roster/other/ known-missing players not already present in that corrected list (see the module docstring's "Extra-players Set ordering" note for this trailing group's order).

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import inject_validated_players
inject_validated_players(["Player One"], box_lineup, ([], []))

is_cached​

is_cached(path: 'str', *, cache_dir: 'Optional[Path]' = None) -> 'bool'

Return whether path already has a cache file on disk.

Parameters

ParameterTypeDefaultDescription
pathstrThe stats.ncaa.org URL path (with its query string, if any).
cache_dirOptional[Path]NoneThe cache root; the active config's cache_dir when None.

Returns

True when cached_path already exists on disk.

is_end_of_game_fouling_vs_fastbreak​

is_end_of_game_fouling_vs_fastbreak(curr_clump: 'ConcurrentClump', event_parser: 'PossessionEvent') -> 'bool'

Check for intentional fouling to prolong the game, specifically so it

can be excluded from being counted as a fast break (is_end_of_game_fouling_vs_fastbreak, LineupUtils.scala:603-656).

Parameters

ParameterTypeDefaultDescription
curr_clumpConcurrentClumpThe clump to classify.
event_parserPossessionEventSelects which side of each event is "attacking".

Returns

True iff the FIRST attacking-side FT-made/FT-missed event in curr_clump.evs is both near the end of a period AND has the attacking team ahead by (0, 10] points; False if no such event exists.

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import is_end_of_game_fouling_vs_fastbreak

is_end_of_game_fouling_vs_fastbreak(curr_clump, event_parser)

is_gen2​

is_gen2(ev: 'RawGameEvent') -> 'bool'

Detect the new/"gen2" NCAA event format (EventUtils.is_gen2,

EventUtils.scala:12-14).

Parameters

ParameterTypeDefaultDescription
evRawGameEventThe raw game event to inspect.

Returns

True if ev.info contains a comma-space (", "), the gen2 format's field separator; False for the old/legacy format.

is_scramble​

is_scramble(curr_clump: 'ConcurrentClump', prev_clumps: 'list[ConcurrentClump]', event_parser: 'PossessionEvent', player_version: 'bool') -> 'tuple[Callable[[RawGameEvent], bool], str]'

Figure out if (each event of) the current clump is part of a

"scramble scenario" following an ORB (is_scramble, LineupUtils .scala:222-597).

Returns a (predicate, debug_tag) tuple -- the tuple shape is load-bearing: the oracle asserts the debug tag string directly ("N/A"/"0a"/"1aa"/"1ab"/"1b"/"2aa"/"2ab").

Parameters

ParameterTypeDefaultDescription
curr_clumpConcurrentClumpThe clump to classify.
prev_clumpslist[ConcurrentClump]Prior merged clumps, most-recent-first.
event_parserPossessionEventSelects which side of each event is "attacking".
player_versionboolUnused -- see the module docstring's is_scramble port notes (the Scala's debug-print gate this flag controls is permanently false regardless of its value).

Returns

(predicate, debug_tag) where predicate(ev) reports whether ev is part of a scramble.

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import is_scramble

predicate, tag = is_scramble(curr_clump, prev_clumps, event_parser, player_version=False)
[predicate(ev) for ev in curr_clump.evs]

is_team_shooting_left_to_start​

is_team_shooting_left_to_start(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]') -> 'tuple[bool, int]'

Infers which side of the SVG court the team under analysis shoots

towards in the first period, from its own made/missed shot locations (ShotEventParser.is_team_shooting_left_to_start, :540-555).

Parameters

ParameterTypeDefaultDescription
sorted_very_raw_eventslist[tuple[int, ShotEvent]]The chronologically-sorted (period, shot) pairs, pre-geometry-transform.

Returns

(team_shooting_left_in_first_period, first_period) -- the first element of sorted_very_raw_events, if any, determines first_period; the majority side (by count) of the team's own (is_off) shots within that period determines the direction.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import is_team_shooting_left_to_start
is_team_shooting_left_to_start([(1, shot_a), (1, shot_b)])

is_women_game​

is_women_game(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]') -> 'bool'

Infers men's vs. women's game from timing evidence

(ShotEventParser.is_women_game, :558-566). Shot-parser-specific variant -- distinct from the play-by-play parser's own is_women_game (Task 5e.3), which uses PbP event timing instead of shot timing; the plan's recon flags both as "its OWN is_women_game variant" per module.

Parameters

ParameterTypeDefaultDescription
sorted_very_raw_eventslist[tuple[int, ShotEvent]]The chronologically-sorted (period, shot) pairs.

Returns

True if at least 4 periods were seen AND no shot was taken with more than 10 minutes showing on the (descending) clock in the very first event (women's quarters are 10 minutes; a shot at >10:00 remaining could only happen in a longer men's period).

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import is_women_game
is_women_game([(1, shot), (2, shot), (3, shot), (4, shot)]) # True

jsoup_text​

jsoup_text(el: 'Optional[Tag]') -> 'str'

JSoup Element.text(): all descendant text, whitespace-collapsed.

JSoup's .text() joins every text node under el (including descendants) and collapses runs of whitespace (spaces, tabs, newlines) into single spaces, trimming the ends. bs4's .get_text() does the joining but not the collapsing, so captured HTML's indentation/newlines would otherwise leak into every extracted value.

Parameters

ParameterTypeDefaultDescription
elOptional[Tag]The element to extract text from, or None.

Returns

The whitespace-collapsed text, or "" if el is None.

Example

from sportsdataverse.mbb.mbb_ncaa_html import jsoup_text, parse_html
soup = parse_html("<td>\n Akin,\tDaniel </td>")
jsoup_text(soup.find("td")) # "Akin, Daniel"

load_proxybonanza_pool​

load_proxybonanza_pool(api_key: 'str', pkg: 'str', *, transport: 'Optional[PoolTransport]' = None) -> "'list[str]'"

Resolve a ProxyBonanza package into a list of http://login:pass@ip:port URLs.

Graduated from dev/ncaa_proxy.py's load_proxy_pool -- same endpoint shape, minus the .Renviron reader (creds are now explicit params, per the creds-hygiene directive).

Endpoint: GET https://api.proxybonanza.com/v1/userpackages/{pkg}.json

Parameters

ParameterTypeDefaultDescription
api_keystrProxyBonanza API key.
pkgstrProxyBonanza package id.
transportOptional[PoolTransport]NoneInjectable (url, headers) -> (status, text) callable for offline testing. Defaults to a curl_cffi GET.

Returns

One http://login:password@ip:port URL per IP in the package.

Example

def fake(url, headers):
return 200, '{"data": {"login": "u", "password": "p", "ippacks": []}}'
pool = load_proxybonanza_pool("key", "pkg", transport=fake)

matching_player​

matching_player(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', tidy_ctx: 'TidyPlayerContext', code_match: 'bool') -> 'bool'

Whether the player in pbp_event matches shot's shooter

(ShotEnrichmentUtils.matching_player, PlayByPlayUtils.scala:638-652).

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched.
pbp_eventMiscGameEventThe candidate play-by-play event.
tidy_ctxTidyPlayerContextThe name-resolution context.
code_matchboolIf True, compare on player code only (looser -- lets a name that resolves to the wrong identity but the right code match); if False, require full ~sportsdataverse.mbb .mbb_ncaa_models.PlayerCodeId equality.

Returns

True if the resolved player matches shot.player under the selected comparison, else False.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import matching_player
matching_player(shot, pbp_event, tidy_ctx, code_match=False)

misspellings​

misspellings(team: 'Optional[TeamId]') -> 'dict[str, str]'

Team-scoped misspelling map, falling back to the generic map

(DataQualityIssues.misspellings, DataQualityIssues.scala:165-322 -- see the module docstring's "fallback semantics" note for why this is a precomputed merge, not a runtime two-level lookup).

Parameters

ParameterTypeDefaultDescription
teamOptional[TeamId]The team to look up team-specific corrections for. None (like any team absent from the table) falls back to generic_misspellings.

Returns

A fresh dict -- the team's misspelling map merged with generic_misspellings, or a copy of generic_misspellings if team has no specific entries.

Example

from sportsdataverse.mbb.mbb_ncaa_data_quality import misspellings
from sportsdataverse.mbb.mbb_ncaa_models import TeamId

misspellings(TeamId("NJIT"))["Lewal, Levi"] # 'Lawal, Levi'
misspellings(TeamId("Some Unlisted Team")) # {} (generic fallback)
misspellings(None) # {} (generic fallback)

ncaa_wbb_box_scores​

ncaa_wbb_box_scores(game_ids: 'Union[str, int, Iterable[Union[str, int]]]', *, multi_games: 'bool' = False, fetcher: 'Optional[Any]' = None, return_as_pandas: 'bool' = False) -> "Union['pl.DataFrame', 'pd.DataFrame']"

Scrape WBB per-player box scores (wbigballR get_box_scores/scrape_box).

Pure delegation to sportsdataverse.mbb.mbb_ncaa_box_stats.ncaa_mbb_box_scores — see it for the column contract, the tolerant header renames, and the fixed multi_games aggregation (R's groups by a Pos column the current markup no longer ships).

Parameters

ParameterTypeDefaultDescription
game_idsUnion[str, int, Iterable[Union[str, int]]]NCAA contest ids; None/NaN entries are dropped.
multi_gamesboolFalseAggregate per player across all games (fixed grouping on player/clean_name/team).
fetcherOptional[Any]NoneOptional injected fetcher exposing fetch_game_individual_stats (tests/offline). Defaults to a fresh NcaaFetcher.with_browser() context per call.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

Per-player box rows (or per-player aggregates with multi_games).

col_nametypedescription
game_idcharacterUnique game identifier.
box_idcharacter
playercharacterPlayer name.
clean_namecharacter
teamcharacterTeam-side label or team identifier.
mpdoubleMinutes played.
ptsdoublePoints scored.
orbdouble
drbdouble
trbdoubleCareer total rebounds.
astdoubleAssists.
todoubleTo.
stldoubleSteals.
blkdoubleBlocks.
fgadoubleField goal attempts.
fgmdoubleField goals made.
fg_pctdoubleField goal percentage (0-1).
tpadouble
tpmdouble
tp_pctdouble
ftadoubleFree throw attempts.
ftmdoubleFree throws made.
ft_pctdoubleFree throw percentage (0-1).
ts_pctdoubleTrue shooting percentage (0-1).
efg_pctdouble
foulsdoublePersonal fouls.
dqdouble
techdouble

Example

from sportsdataverse.wbb.wbb_ncaa_box_stats import ncaa_wbb_box_scores
box = ncaa_wbb_box_scores(["5722355"])
print(box.shape)

ncaa_wbb_date_games​

ncaa_wbb_date_games(date: 'Optional[str]' = None, *, conference: 'str' = 'All', conference_id: 'Optional[int]' = None, fetcher: "Optional['NcaaFetcher']" = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Discover every NCAA WBB game played on a date (wbigballR get_date_games).

Same engine as sportsdataverse.mbb.mbb_ncaa_scoreboard.ncaa_mbb_date_games with the WBB season_divisions table bound (see the module docstring for the 2010-11..2025-26 coverage caveat).

Parameters

ParameterTypeDefaultDescription
dateOptional[str]None"MM/DD/YYYY". Defaults to yesterday (R default).
conferencestr'All'Conference name filter (e.g. "SEC", "Summit League"); case/punctuation-insensitive. Default "All" (every conference). Unknown names raise.
conference_idOptional[int]NoneExplicit stats.ncaa.org conference id; overrides conference when given.
fetcherOptional['NcaaFetcher']NoneInjectable ~sportsdataverse.mbb.mbb_ncaa_fetch. NcaaFetcher (tests pass an offline fake). None uses NcaaFetcher.with_browser().
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per game — see the MBB sibling for the full SCOREBOARD_SCHEMA column contract.

col_nametypedescription
datecharacterDate in YYYY-MM-DD format.
start_timecharacter
homecharacterHome.
awaycharacterAway record.
box_idcharacter
game_idcharacterUnique game identifier.
home_scorecharacterHome team score at the time of the play.
away_scorecharacterAway team score at the time of the play.
attendancecharacterReported attendance.
neutral_sitelogicalNeutral site.
home_winsintegerHome team's wins.
home_lossesintegerHome team's losses.
away_winsintegerAway team's wins.
away_lossesintegerAway team's losses.

Example

from sportsdataverse.wbb.wbb_ncaa_scoreboard import ncaa_wbb_date_games
games = ncaa_wbb_date_games("12/05/2024")
print(games.shape)

ncaa_wbb_join_pbp_shots​

ncaa_wbb_join_pbp_shots(pbp: 'pl.DataFrame', shots: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Attach WBB chart shots onto the pbp frame (pure delegation).

See sportsdataverse.mbb.mbb_ncaa_shots.ncaa_mbb_join_pbp_shots for the matching rules (FG-only, within-second same-result sequence) and the joined 40-column contract.

Parameters

ParameterTypeDefaultDescription
pbpDataFrame35-column snake_case pbp frame (ncaa_wbb_play_by_play).
shotsDataFrameShots frame from ncaa_wbb_shot_locations.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

The pbp frame with shot columns attached (unmatched rows NA-filled).

col_nametypedescription
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
awaycharacterAway record.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
clockcharacterGame clock value.
game_timecharacterGame start time.
game_secondsinteger
home_scoreintegerHome team score at the time of the play.
away_scoreintegerAway team score at the time of the play.
event_teamcharacter
event_descriptioncharacter
player_1character
player_2character
event_typecharacterEvent / play type code (V2 PBP).
event_resultcharacter
shot_valueintegerPoint value of the shot (2 or 3).
event_lengthinteger
poss_numinteger
poss_teamcharacter
poss_lengthinteger
is_transitionlogical
home_1character
home_2character
home_3character
home_4character
home_5character
away_1character
away_2character
away_3character
away_4character
away_5character
statuscharacterStatus label.
is_garbage_timelogical
sub_deviateinteger
teamcharacterTeam-side label or team identifier.
playercharacterPlayer name.
xdoubleX.
ydoubleY.
shot_distdouble

Example

from sportsdataverse.wbb.wbb_ncaa_shots import ncaa_wbb_join_pbp_shots
joined = ncaa_wbb_join_pbp_shots(pbp, shots)
print(joined.shape)

ncaa_wbb_player_stats​

ncaa_wbb_player_stats(pbp: 'pl.DataFrame', *, multi_games: 'bool' = False, simple: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Aggregate WBB play-by-play into per-player box stats (wbigballR get_player_stats).

Pure delegation to sportsdataverse.mbb.mbb_ncaa_stats_agg.ncaa_mbb_player_stats — see it for the algorithm and column contracts.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (ncaa_wbb_game_pbp output).
multi_gamesboolFalseAggregate across games per (player, team) — the season-stat surface.
simpleboolFalseReturn the reduced surface without the transition / assisted / putback / block-location splits.
fix_tip_inboolTrueCount the real "Tip In" vocabulary (default); False reproduces R's "Tip-In" bug for oracle parity.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per player+team (+game when multi_games=False).

col_nametypedescription
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
awaycharacterAway record.
teamcharacterTeam-side label or team identifier.
playercharacterPlayer name.
minsdouble
o_possdouble
ptsdoublePoints scored.
orbdouble
drbdouble
astdoubleAssists.
stldoubleSteals.
blkdoubleBlocks.
tovdoubleTurnovers.
pfdoublePersonal fouls.
ts_pctdoubleTrue shooting percentage (0-1).
efg_pctdouble
fgmdoubleField goals made.
fgadoubleField goal attempts.
fg_pctdoubleField goal percentage (0-1).
tpmdouble
tpadouble
tp_pctdouble
ftmdoubleFree throws made.
ftadoubleFree throw attempts.
ft_pctdoubleFree throw percentage (0-1).
rimmdouble
rimadouble
rim_pctdouble
midmdouble
midadouble
mid_pctdouble
pbackmdouble
pbackadouble
pback_pctdouble
blk_rimdouble
blk_middouble
blk_threedouble
pct_fga_transdouble
pct_tpa_transdouble
pct_rima_transdouble
pct_fgm_transdouble
pct_tpm_transdouble
pct_rimm_transdouble
pct_fgm_astdouble
pct_tpm_astdouble
pct_rimm_astdouble
pts_transdouble
orb_transdouble
drb_transdouble
ast_transdouble
stl_transdouble
blk_transdouble
tov_transdouble
ts_pct_transdouble
efg_pct_transdouble
fgm_transdouble
fga_transdouble
fg_pct_transdouble
tpm_transdouble
tpa_transdouble
tp_pct_transdouble
ftm_transdouble
fta_transdouble
ft_pct_transdouble
rimm_transdouble
rima_transdouble
rim_pct_transdouble
midm_transdouble
mida_transdouble
mid_pct_transdouble
pts_halfdouble
orb_halfdouble
drb_halfdouble
ast_halfdouble
stl_halfdouble
blk_halfdouble
tov_halfdouble
ts_pct_halfdouble
efg_pct_halfdouble
fgm_halfdouble
fga_halfdouble
fg_pct_halfdouble
tpm_halfdouble
tpa_halfdouble
tp_pct_halfdouble
ftm_halfdouble
fta_halfdouble
ft_pct_halfdouble
rimm_halfdouble
rima_halfdouble
rim_pct_halfdouble
midm_halfdouble
mida_halfdouble
mid_pct_halfdouble
pts_astdouble
fgm_astdouble
tpm_astdouble
rimm_astdouble
midm_astdouble
pts_unastdouble
efg_pct_unastdouble
fgm_unastdouble
fga_unastdouble
fg_pct_unastdouble
tpm_unastdouble
tpa_unastdouble
tp_pct_unastdouble
rimm_unastdouble
rima_unastdouble
rim_pct_unastdouble
midm_unastdouble
mida_unastdouble
mid_pct_unastdouble

Example

from sportsdataverse.wbb.wbb_ncaa_stats_agg import ncaa_wbb_player_stats
stats = ncaa_wbb_player_stats(pbp)
print(stats.shape)

ncaa_wbb_possessions​

ncaa_wbb_possessions(pbp: 'pl.DataFrame', *, simple: 'bool' = False, fix_cross_game_leak: 'bool' = True, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Aggregate WBB play-by-play into one row per possession (wbigballR get_possessions).

Pure delegation to sportsdataverse.mbb.mbb_ncaa_possession_seg.ncaa_mbb_possessions — see it for the algorithm, the 28/17-column contracts, and the fixed-vs- faithful flag convention.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (ncaa_wbb_game_pbp output).
simpleboolFalseReturn only the 17-column possession/points frame.
fix_cross_game_leakboolTrueWhen True (default, and the CORRECT behavior), window the start_event_type lag with .over("game_id") so a game's first possession does not inherit the previous game's last event. When False, reproduce R's ungrouped dplyr::lag (all_functions.R:3698). Parity tests pass False.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per possession.

col_nametypedescription
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
awaycharacterAway record.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
poss_numinteger
poss_teamcharacter
home_1character
home_2character
home_3character
home_4character
home_5character
away_1character
away_2character
away_3character
away_4character
away_5character
home_scoreintegerHome team score at the time of the play.
away_scoreintegerAway team score at the time of the play.
ptsintegerPoints scored.
is_assistedinteger
is_transitioninteger
is_garbage_timeinteger
start_event_typecharacter
first_shot_timeinteger
first_shot_typecharacter
last_event_timeinteger
last_event_typecharacter

Example

from sportsdataverse.wbb.wbb_ncaa_possession_seg import ncaa_wbb_possessions
poss = ncaa_wbb_possessions(pbp)
print(poss.shape)

# Faithful (R-buggy) start-event lag

poss = ncaa_wbb_possessions(pbp, fix_cross_game_leak=False)

ncaa_wbb_shot_locations​

ncaa_wbb_shot_locations(game_ids: 'Sequence[object]', *, fetcher: 'Optional[Any]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Scrape WBB shot locations for one or more games.

Same driver as sportsdataverse.mbb.mbb_ncaa_shots.ncaa_mbb_shot_locations with the quarters period_model bound — see the mbb sibling for the parse algorithm and the ~sportsdataverse.mbb.mbb_ncaa_shots.SHOTS_SCHEMA contract.

Parameters

ParameterTypeDefaultDescription
game_idsSequence[object]NCAA contest ids; None/NaN entries are dropped.
fetcherOptional[Any]NoneOptional injected fetcher exposing fetch_game_box (tests/offline). Defaults to a fresh NcaaFetcher.with_browser() context per call.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

All games' shots row-bound (zero-row schema frame when none found).

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
clockcharacterGame clock value.
game_secondsinteger
teamcharacterTeam-side label or team identifier.
playercharacterPlayer name.
shot_resultcharacterShot result ('Made' / 'Missed').
xdoubleX.
ydoubleY.
shot_distdouble

Example

from sportsdataverse.wbb.wbb_ncaa_shots import ncaa_wbb_shot_locations
shots = ncaa_wbb_shot_locations(["5722355"])
print(shots.shape)

ncaa_wbb_team_roster​

ncaa_wbb_team_roster(team_id: 'Optional[int]' = None, *, team: 'Optional[str]' = None, season: 'Optional[str]' = None, fetcher: "Optional['NcaaFetcher']" = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Scrape a women's team roster from stats.ncaa.org.

Port of wbigballR get_team_roster with name resolution fixed to the WBB crosswalk (see the module docstring). The roster parser itself is league-agnostic; algorithm detail: sportsdataverse.mbb.mbb_ncaa_schedule.ncaa_mbb_team_roster.

Parameters

ParameterTypeDefaultDescription
team_idOptional[int]Nonestats.ncaa.org team id (changes every season).
teamOptional[str]NoneSchool name, e.g. "South Carolina".
seasonOptional[str]NoneSeason string, e.g. "2024-25"; required with team.
fetcherOptional['NcaaFetcher']NoneInjectable ~sportsdataverse.mbb.mbb_ncaa_fetch. NcaaFetcher; defaults to a fresh browser-transport fetcher.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per player — see ~sportsdataverse.mbb.mbb_ncaa_schedule.parse_ncaa_bb_team_roster for the column contract.

col_nametypedescription
gpcharacterGames played.
gscharacterGames started.
jerseycharacterJersey number worn by the player.
namecharacterDisplay name.
classcharacterCollege class / draft eligibility note.
positioncharacterListed roster position (G, F, C, etc.).
heightcharacterPlayer height (string e.g. '6-2' or inches).
hometowncharacterPlayer hometown.
high_schoolcharacter
playercharacterPlayer name.
clean_namecharacter
ht_inchesinteger
player_idcharacterUnique player identifier.

Example

from sportsdataverse.wbb.wbb_ncaa_schedule import ncaa_wbb_team_roster
df = ncaa_wbb_team_roster(team="South Carolina", season="2024-25")
print(df.select("jersey", "player", "ht_inches").head())

ncaa_wbb_team_schedule​

ncaa_wbb_team_schedule(team_id: 'Optional[int]' = None, *, team: 'Optional[str]' = None, season: 'Optional[str]' = None, fetcher: "Optional['NcaaFetcher']" = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Scrape a women's team's season schedule from stats.ncaa.org.

Port of wbigballR get_team_schedule with name resolution fixed to the WBB crosswalk (see the module docstring). Algorithm detail: sportsdataverse.mbb.mbb_ncaa_schedule.ncaa_mbb_team_schedule.

Parameters

ParameterTypeDefaultDescription
team_idOptional[int]Nonestats.ncaa.org team id (changes every season).
teamOptional[str]NoneSchool name, e.g. "South Carolina".
seasonOptional[str]NoneSeason string, e.g. "2024-25"; required with team.
fetcherOptional['NcaaFetcher']NoneInjectable ~sportsdataverse.mbb.mbb_ncaa_fetch. NcaaFetcher; defaults to a fresh browser-transport fetcher.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per scheduled game — see ~sportsdataverse.mbb.mbb_ncaa_schedule.parse_ncaa_bb_team_schedule for the column contract.

col_nametypedescription
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
home_scoreintegerHome team score at the time of the play.
awaycharacterAway record.
away_scoreintegerAway team score at the time of the play.
box_idcharacter
game_idcharacterUnique game identifier.
is_neutrallogical
detailcharacterDetail.
attendanceintegerReported attendance.

Example

from sportsdataverse.wbb.wbb_ncaa_schedule import ncaa_wbb_team_schedule
df = ncaa_wbb_team_schedule(team="South Carolina", season="2024-25")
print(df.shape)

ncaa_wbb_team_stats​

ncaa_wbb_team_stats(pbp: 'pl.DataFrame', *, include_transition: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Aggregate WBB play-by-play into per-team game stats (wbigballR get_team_stats).

Pure delegation to sportsdataverse.mbb.mbb_ncaa_stats_agg.ncaa_mbb_team_stats — see it for the algorithm and column contract.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (ncaa_wbb_game_pbp output).
include_transitionboolFalseAppend the trans/half split surface.
fix_tip_inboolTrueCount the real "Tip In" vocabulary (default); False reproduces R's "Tip-In" bug for oracle parity.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per team per game.

col_nametypedescription
game_idcharacterUnique game identifier.
homecharacterHome.
awaycharacterAway record.
teamcharacterTeam-side label or team identifier.
minsdouble
o_minsdouble
d_minsdouble
o_possdouble
d_possdouble
ortgdouble
drtgdouble
netrtgdouble
ptsdoublePoints scored.
d_ptsdouble
fgadoubleField goal attempts.
d_fgadouble
fgmdoubleField goals made.
d_fgmdouble
tpadouble
d_tpadouble
tpmdouble
d_tpmdouble
ftadoubleFree throw attempts.
d_ftadouble
ftmdoubleFree throws made.
d_ftmdouble
rimadouble
d_rimadouble
rimmdouble
d_rimmdouble
orbdouble
d_orbdouble
drbdouble
d_drbdouble
blkdoubleBlocks.
d_blkdouble
todoubleTo.
d_todouble
astdoubleAssists.
d_astdouble
e_possdouble
fg_pctdoubleField goal percentage (0-1).
d_fg_pctdouble
tppdouble
d_tppdouble
ftpdouble
d_ftpdouble
efg_pctdouble
d_efg_pctdouble
ts_pctdoubleTrue shooting percentage (0-1).
d_ts_pctdouble
rim_pctdouble
d_rim_pctdouble
mid_pctdouble
d_mid_pctdouble
tp_ratedouble
d_tp_ratedouble
rim_ratedouble
d_rim_ratedouble
mid_ratedouble
d_mid_ratedouble
ft_ratedoubleFt rate.
d_ft_ratedouble
ast_ratedouble
d_ast_ratedouble
to_ratedoubleTo rate.
d_to_ratedouble
blk_ratedouble
o_blk_ratedouble
orb_pctdoubleOffensive rebound percentage.
drb_pctdoubleDefensive rebound percentage.
time_per_possdouble
d_time_per_possdouble

Example

from sportsdataverse.wbb.wbb_ncaa_stats_agg import ncaa_wbb_team_stats
team = ncaa_wbb_team_stats(pbp)
print(team.shape)

phase1_shot_event_enrichment​

phase1_shot_event_enrichment(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]', second_half_override: 'Optional[set[int]]' = None) -> 'list[ShotEvent]'

The court-geometry enrichment pass: ascending time, coordinate

transform + geo synthesis, and the self-correcting side-flip re-run (ShotEventParser.phase1_shot_event_enrichment, :415-528).

For each shot: compute the ascending game time, decide (from is_team_shooting_left_to_start + which half the period falls in) whether the shot's side needs flipping, run transform_shot_location to get both the believed-correct and alternative (mirrored) locations, keep whichever is closer to the basket (a >1.2x distance advantage for the "alternative" wins, or ANY shot taken with <0.1 min left on the clock always keeps the original -- a half-court heave near the buzzer is plausible, so the tie-break favors trusting the raw geometry there), then synthesize a lat/lon.

After all shots are processed, if any period had >=6 shots AND more than 75% of them came back implausibly long-distance (>50ft), the whole pass re-runs ONCE with those periods' orientation flipped (the self-correcting part) -- second_half_override is None on the initial call and a non-None set on the one allowed retry, preventing infinite recursion.

Parameters

ParameterTypeDefaultDescription
sorted_very_raw_eventslist[tuple[int, ShotEvent]]The chronologically-sorted (period, shot) pairs from parse_shot_html, pre-geometry-transform.
second_half_overrideOptional[set[int]]NoneThe set of periods whose second_half_switch orientation should be inverted (the self-correction re-run's input); None on the first call.

Returns

The fully court-geometry-enriched shots, in the same order as sorted_very_raw_events.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import phase1_shot_event_enrichment
shots = phase1_shot_event_enrichment([(1, very_raw_shot)])