MBB — additional Python functions — stats.ncaa.org: ncaa_mbb–validate_lineup
ncaa_mbb_shot_locations
ncaa_mbb_shot_locations(game_ids: "'Sequence[object]'", *, fetcher: 'Optional[_SupportsFetchGameBox]' = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, Any]'"
Scrape MBB shot locations for one or more games (bigballR
get_shot_locations, get_shot_locations.R:3-89).
Fetches each game's stats.ncaa.org/contests/{id}/box_score page and
parses the embedded shot-chart JS through parse_ncaa_bb_shots.
NA ids are dropped up front (R :5); per-game "shots found" messages
go to the module logger (R message, :69-70).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
game_ids | Sequence[object] | NCAA contest ids; None/NaN entries are dropped. | |
fetcher | Optional[_SupportsFetchGameBox] | None | Optional injected fetcher exposing fetch_game_box (for tests/offline use). Defaults to a fresh NcaaFetcher.with_browser() context per call. |
return_as_pandas | bool | False | Return a pandas DataFrame instead of polars. |
Returns
All games' shots row-bound (zero-row SHOTS_SCHEMA frame when no ids survive or no charts are found).
| col_name | type | description |
|---|---|---|
game_id | character | Unique game identifier. |
period | integer | Period of the game (1-4 quarters; 5+ for OT). |
clock | character | Game clock value. |
game_seconds | integer | |
team | character | Team-side label or team identifier. |
player | character | Player name. |
shot_result | character | Shot result ('Made' / 'Missed'). |
x | double | X. |
y | double | Y. |
shot_dist | double |
Example
from sportsdataverse.mbb.mbb_ncaa_shots import ncaa_mbb_shot_locations
df = ncaa_mbb_shot_locations(["6470186", "6479639"])
print(df.shape)
# Offline with an injected fetcher
df = ncaa_mbb_shot_locations(["6470186"], fetcher=my_fetcher)
# Pipeline next step (one line)
df.group_by("team").agg(pl.col("shot_dist").mean()).head()
ncaa_mbb_team_stats
ncaa_mbb_team_stats(pbp: 'pl.DataFrame', *, include_transition: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'
Aggregate bigballR-contract play-by-play into per-team game stats.
Port of bigballR get_team_stats (all_functions.R:2530-2538): the
ten on-court columns are blanked so every row shares one "lineup", then
get_lineups (ncaa_mbb_lineups) runs per game and the lineup
key columns are dropped — yielding two rows (one per team) per game.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
pbp | DataFrame | Play-by-play frame in the sdv-py 35-column snake_case bigballR contract. May span multiple games. | |
include_transition | bool | False | When True, append the trans/half split surface plus o_trans_pct/d_trans_pct. |
fix_tip_in | bool | True | When True (default), rim stats count the scrape engine's real "Tip In" vocabulary; False reproduces R's literal "Tip-In" bug for oracle parity. |
return_as_pandas | bool | False | Return a pandas.DataFrame instead of polars. |
Returns
pl.DataFrame (or pd.DataFrame) with one row per team per game — TEAM_STATS_COLUMNS (73) or TEAM_STATS_TRANSITION_COLUMNS with include_transition=True. Games ordered by the Utf8 game_id byte sort (R's do() sorts a numeric ID — identical for equal-width ids), teams within a game byte-sorted. Empty input yields an empty frame with the documented schema.
| col_name | type | description |
|---|---|---|
game_id | character | Unique game identifier. |
home | character | Home. |
away | character | Away record. |
team | character | Team-side label or team identifier. |
mins | double | |
o_mins | double | |
d_mins | double | |
o_poss | double | |
d_poss | double | |
ortg | double | |
drtg | double | |
netrtg | double | |
pts | double | Points scored. |
d_pts | double | |
fga | double | Field goal attempts. |
d_fga | double | |
fgm | double | Field goals made. |
d_fgm | double | |
tpa | double | |
d_tpa | double | |
tpm | double | |
d_tpm | double | |
fta | double | Free throw attempts. |
d_fta | double | |
ftm | double | Free throws made. |
d_ftm | double | |
rima | double | |
d_rima | double | |
rimm | double | |
d_rimm | double | |
orb | double | |
d_orb | double | |
drb | double | |
d_drb | double | |
blk | double | Blocks. |
d_blk | double | |
to | double | To. |
d_to | double | |
ast | double | Assists. |
d_ast | double | |
e_poss | double | |
fg_pct | double | Field goal percentage (0-1). |
d_fg_pct | double | |
tpp | double | |
d_tpp | double | |
ftp | double | |
d_ftp | double | |
efg_pct | double | |
d_efg_pct | double | |
ts_pct | double | True shooting percentage (0-1). |
d_ts_pct | double | |
rim_pct | double | |
d_rim_pct | double | |
mid_pct | double | |
d_mid_pct | double | |
tp_rate | double | |
d_tp_rate | double | |
rim_rate | double | |
d_rim_rate | double | |
mid_rate | double | |
d_mid_rate | double | |
ft_rate | double | Ft rate. |
d_ft_rate | double | |
ast_rate | double | |
d_ast_rate | double | |
to_rate | double | To rate. |
d_to_rate | double | |
blk_rate | double | |
o_blk_rate | double | |
orb_pct | double | Offensive rebound percentage. |
drb_pct | double | Defensive rebound percentage. |
time_per_poss | double | |
d_time_per_poss | double |
Example
from sportsdataverse.mbb.mbb_ncaa_stats_agg import ncaa_mbb_team_stats
teams = ncaa_mbb_team_stats(pbp)
print(teams.shape)
# Transition splits, pandas out
df_pd = ncaa_mbb_team_stats(pbp, include_transition=True, return_as_pandas=True)
# Pipeline next step (one line)
teams.sort("netrtg", descending=True).head()
phase1_shot_event_enrichment
phase1_shot_event_enrichment(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]', second_half_override: 'Optional[set[int]]' = None) -> 'list[ShotEvent]'
The court-geometry enrichment pass: ascending time, coordinate
transform + geo synthesis, and the self-correcting side-flip re-run
(ShotEventParser.phase1_shot_event_enrichment, :415-528).
For each shot: compute the ascending game time, decide (from
is_team_shooting_left_to_start + which half the period falls in)
whether the shot's side needs flipping, run transform_shot_location
to get both the believed-correct and alternative (mirrored) locations,
keep whichever is closer to the basket (a >1.2x distance advantage for
the "alternative" wins, or ANY shot taken with <0.1 min left on the
clock always keeps the original -- a half-court heave near the buzzer
is plausible, so the tie-break favors trusting the raw geometry there),
then synthesize a lat/lon.
After all shots are processed, if any period had >=6 shots AND more than
75% of them came back implausibly long-distance (>50ft), the whole pass
re-runs ONCE with those periods' orientation flipped (the self-correcting
part) -- second_half_override is None on the initial call and a
non-None set on the one allowed retry, preventing infinite recursion.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
sorted_very_raw_events | list[tuple[int, ShotEvent]] | The chronologically-sorted (period, shot) pairs from parse_shot_html, pre-geometry-transform. | |
second_half_override | Optional[set[int]] | None | The set of periods whose second_half_switch orientation should be inverted (the self-correction re-run's input); None on the first call. |
Returns
The fully court-geometry-enriched shots, in the same order as sorted_very_raw_events.
Example
from sportsdataverse.mbb.mbb_ncaa_shot_parser import phase1_shot_event_enrichment
shots = phase1_shot_event_enrichment([(1, very_raw_shot)])
playwright_transport
playwright_transport(*, headless_new: 'bool' = True, challenge_wait_ms: 'int' = 8000, nav_timeout_ms: 'int' = 45000, user_agent: 'Optional[str]' = None, solve_attempts: 'int' = 3, relaunch_backoff: 'float' = 2.0) -> "'_PlaywrightTransport'"
Build the suggested stats.ncaa.org game-detail scraping transport.
Drives a real Chromium via Playwright in Chrome's new-headless mode
(--headless=new) to clear the Akamai bm-verify challenge that
curl_cffi cannot, then serves raw server HTML for the 5a-5e parsers.
Playwright is a lazy optional import (not a hard dependency); a clear
ImportError fires on first use if it is missing.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
headless_new | bool | True | Use --headless=new (real-GPU render, no window) -- the default and the proven-working mode. False runs old headless (headless_shell), which Akamai flags -- avoid. |
challenge_wait_ms | int | 8000 | Milliseconds to let the bm-verify sensor run after the first navigation. |
nav_timeout_ms | int | 45000 | Per-navigation timeout. |
user_agent | Optional[str] | None | Override the Chrome UA string. |
solve_attempts | int | 3 | |
relaunch_backoff | float | 2.0 |
Returns
A stateful, callable FetchTransport reusing one browser for the session. Close it when done (it is a context manager, has close(), and registers an atexit safety net).
Example
from sportsdataverse.mbb.mbb_ncaa_fetch import NcaaFetcher
with NcaaFetcher.with_browser() as fetcher:
pbp = fetcher.fetch_game_pbp("1613299") # raw PBP HTML
box = fetcher.fetch_game_individual_stats("1613299") # raw box HTML
# -> feed to get_box_lineup / create_lineup_data (mbb_ncaa_*_parser)
remove_diacritics
remove_diacritics(fragment: 'str') -> 'str'
Strip diacritical marks, e.g. "Juhász" -> "Juhasz"
(ExtractorUtils.scala:38-43: NFD normalization then removal of the
combining-diacritical-marks block).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
fragment | str | Any string (a full player name or a name fragment). |
Returns
The string with combining marks removed.
Example
from sportsdataverse.mbb.mbb_ncaa_stints import remove_diacritics
print(remove_diacritics("Dorka Juhász")) # "Dorka Juhasz"
remove_html_encoding
remove_html_encoding(html_str: 'str') -> 'str'
Undo a handful of literal HTML entity escapes (``ExtractorUtils
.remove_html_encoding, ExtractorUtils.scala:25-33). **Scope addition, Task 5e.5** -- the first consumer is mbb_ncaa_shot_parser.parse_shot_html(theplayername / shooting team name extracted from an SVG shot's
In practice bs4/lxml already decode standard HTML entities (',
", `&``) while parsing text nodes, so this is usually a
no-op by the time it runs on already-parsed text -- ported anyway for
exact behavioral parity with any double-escaped input the upstream
Scala guards against (JSoup has the same auto-decoding behavior, so the
Scala original is equally a defensive no-op in the common case).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
html_str | str | Any string, typically already-parsed element text. |
Returns
html_str with '/"/& replaced by their literal characters, only if "&" appears at all (short-circuit matching the Scala's if (html_str.indexOf("&") >= 0) guard).
Example
from sportsdataverse.mbb.mbb_ncaa_stints import remove_html_encoding
remove_html_encoding("De'Shayne") # "De'Shayne"
remove_html_encoding("Plain Name") # "Plain Name" (unchanged)
reorder_and_reverse
reorder_and_reverse(reversed_partial_events: 'Iterable[PlayByPlayEvent]') -> 'list[PlayByPlayEvent]'
Orders same-minute play-by-play events so subs never enclose the plays
they logically precede/follow (ExtractorUtils.scala:435-599).
Groups consecutive events sharing the same min into a block (the
input arrives in descending/reverse-chronological order, so blocks are
discovered and internally accumulated in reverse too), then -- for any
block containing a sub -- reorders it via inner_sort: events
referencing a subbed-OUT player (or scoring no higher than the sub) land
in a pre-sub group, the subs themselves come next (in ascending-score
order), and events referencing a subbed-IN player (or scoring higher
than the sub) land in a trailing post-sub group. Free-throw attempts
sharing the sub's inferred "direction" (team vs. opponent, inferred from
the nearest preceding shot/FT/foul) are pulled into the pre-sub group
unless the shooter is one of the players being subbed in. Blocks with no
sub are returned unchanged apart from the initial score-based sort.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
reversed_partial_events | Iterable[PlayByPlayEvent] | Events for one lineup event, in reverse-chronological (descending-time) order -- the natural order encountered walking play-by-play text bottom-up. |
Returns
The same events, forward-chronological (ascending time), with each same-minute block internally reordered so no sub encloses a play it logically shouldn't.
Example
from sportsdataverse.mbb.mbb_ncaa_models import Score
from sportsdataverse.mbb.mbb_ncaa_stints import (
OtherTeamEvent,
SubInEvent,
reorder_and_reverse,
)
events = [
SubInEvent(0.4, Score(0, 0), "player1"),
OtherTeamEvent(0.4, Score(0, 0), "rebound"),
]
reorder_and_reverse(events)
# [OtherTeamEvent(...), SubInEvent(...)]
reset_config
reset_config() -> 'NcaaFetchConfig'
Reset the active config to its env-var-derived defaults.
Returns
The live singleton, now holding the env-var-derived defaults again.
Example
from sportsdataverse.mbb.mbb_ncaa_fetch import update_config, reset_config
update_config(timeout=5)
reset_config()
right_kind_of_shot
right_kind_of_shot(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', strict: 'bool') -> 'bool'
Whether pbp_event's shot type is compatible with shot's
distance and make/miss (ShotEnrichmentUtils.right_kind_of_shot,
PlayByPlayUtils.scala:659-679).
The distance-in-the-data is approximate, so exact 2-vs-3 discrimination is impossible; this only rules out the obvious mismatches (a clearly-short shot matched to a 3, or vice versa) and always requires make/miss agreement.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
shot | ShotEvent | The shot being enriched (pts/dist read). | |
pbp_event | MiscGameEvent | The candidate play-by-play event. | |
strict | bool | If True, also apply the distance gate; if False, only the make/miss agreement is required. |
Returns
True if the event could plausibly be this shot.
Example
from sportsdataverse.mbb.mbb_ncaa_pbp_glue import right_kind_of_shot
right_kind_of_shot(shot, pbp_event, strict=True)
run_iterative_adjustment_with_hca
run_iterative_adjustment_with_hca(teams: 'Sequence[TeamDetail]', team_by_name: 'dict[str, TeamDetail]', fields: 'Sequence[str]', league_averages: 'LeagueAverages', poss_splits: 'dict[str, PossessionSplits]', *, max_iterations: 'int' = 100, tolerance: 'float' = 1e-06) -> 'IterationResult'
KenPom-style SoS + HCA fixed-point solver (runIterativeAdjustmentWithHCA, ts:306-527).
Each iteration (Jacobi -- all teams read the previous iteration's adjustments, then commit together):
- Per team/field, adjust every game
adj_game = raw_game * (league / (opp_adj +/- hca))and take the weighted mean; a field with no valid games keeps its current value. - Re-estimate per-field HCA from home/away possession-imbalance residuals
hca = sum((raw - pred) * |imbalance|) / sum(|imbalance|)over teams with|imbalance| >= IMBALANCE_MIN.
Stops when the max per-team/field change drops below tolerance or after
max_iterations sweeps (the HCA re-estimate still runs on the final
sweep). The cross-guard on the per-game branch, the asymmetric residual
prediction, and the cross-named opponent strengths are all preserved -- see
the module docstring's landmine list.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
teams | Sequence[TeamDetail] | The teams to solve over. | |
team_by_name | dict[str, TeamDetail] | {team_name: team_detail} for opponent lookup. | |
fields | Sequence[str] | The stat fields to solve. | |
league_averages | LeagueAverages | Output of compute_league_averages_from_per_game. | |
poss_splits | dict[str, PossessionSplits] | {team_name: PossessionSplits }. | |
max_iterations | int | 100 | Iteration cap (default MAX_ITERATIONS; pin to 1 to inspect a single sweep). |
tolerance | float | 1e-06 | Convergence tolerance (default TOLERANCE). |
Returns
An IterationResult (adj_values, hca_per_field).
Example
from sportsdataverse.mbb.mbb_ncaa_strength import (
STRENGTH_ADJUSTED_FIELDS,
compute_league_averages_from_per_game,
compute_possession_splits,
run_iterative_adjustment_with_hca,
)
by_name = {t["team_name"]: t for t in teams}
league = compute_league_averages_from_per_game(teams)
splits = {t["team_name"]: compute_possession_splits(t) for t in teams}
result = run_iterative_adjustment_with_hca(
teams, by_name, STRENGTH_ADJUSTED_FIELDS, league, splits,
)
print(result.hca_per_field["3p"]["hca_off"])
same_school
same_school(a: 'str', b: 'str') -> 'bool'
Whether two team-name spellings denote the same school.
Exact match, or both names inside one team_name_equivalents class.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
a | str | One team-name spelling. | |
b | str | The other team-name spelling. |
Returns
True if the two names refer to the same school.
Example
from sportsdataverse.mbb.mbb_ncaa_data_quality import same_school
same_school("New Orleans", "LSU New Orleans") # True
same_school("Miami (FL)", "Miami (OH)") # False
select_contains
select_contains(root: 'Tag', selector: 'str', text: 'str') -> 'list[Tag]'
JSoup root.select(sel + ":contains(text)"): candidates whose full
text (own + every descendant's) case-insensitively CONTAINS text as
a plain substring -- not a regex (Task 5e.2 addition; see the module
docstring's "Critical divergence" note).
JSoup's :contains() is documented case-insensitive substring
containment; soupsieve's :-soup-contains() (the non-deprecated
spelling of its :contains()) is case-SENSITIVE, with no
case-insensitive variant of its own. Reproducing JSoup's actual
semantics therefore needs this helper rather than :-soup-contains().
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
root | Tag | The element to search within. | |
selector | str | A plain (soupsieve-legal) CSS selector for the structural part of the match (everything before :contains). | |
text | str | The plain substring each candidate's collapsed text must case-insensitively contain. |
Returns
Every selector match whose jsoup_text case-insensitively contains text, in document order.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_contains
soup = parse_html("<td>game date:</td><td>Location:</td>")
select_contains(soup, "td", "Game Date:") # [<td>game date:</td>]
select_matching
select_matching(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'
JSoup root.select(sel + ":matches(regex)"): candidates whose full
text (own + every descendant's) matches regex.
Soupsieve has no :matches() pseudo-class equivalent, so this runs the
plain structural selector first, then filters by re.search
over each candidate's jsoup_text (own text plus descendants',
matching JSoup's :matches() semantics -- as opposed to
select_matching_own's own-text-only :matchesOwn()).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
root | Tag | The element to search within. | |
selector | str | A plain (soupsieve-legal) CSS selector. | |
regex | str | The pattern each candidate's collapsed text must re.search-match. |
Returns
Every selector match whose jsoup_text contains a regex match, in document order.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching
soup = parse_html("<div><p>Home Team</p><p>Away Team</p></div>")
select_matching(soup, "p", r"^Home") # [<p>Home Team</p>]
select_matching_own
select_matching_own(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'
JSoup root.select(sel + ":matchesOwn(regex)"): candidates whose
OWN text only (excluding descendant elements' text) matches regex.
JSoup's Element.ownText() walks only the element's direct
TextNode children, not text nested inside child elements -- the
same distinction bs4 draws between a tag's direct
bs4.NavigableString children and its full .get_text().
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
root | Tag | The element to search within. | |
selector | str | A plain (soupsieve-legal) CSS selector. | |
regex | str | The pattern each candidate's own (whitespace-collapsed) text must re.search-match. |
Returns
Every selector match whose own text contains a regex match, in document order.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching_own
soup = parse_html('<div class="card-header">Coach <b>Info</b></div>')
select_matching_own(soup, "div.card-header", r"^Coach")
# [<div class="card-header">Coach <b>Info</b></div>]
shot_js_to_html
shot_js_to_html(js: 'str') -> 'list[Tag]'
Converts client-side addShot(...) JS calls into parseable
circle.shot HTML, for pages where the shot map is built on the fly
rather than baked into the initial HTML (ShotEventParser .shot_js_to_html, :266-283). See the module docstring's "Scala
idiom decision" note -- the Scala's builders/browser parameters
are dropped here since the Scala body never actually uses them.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
js | str | The concatenated <script> text containing one or more addShot(x, y, ..., 'title', ...) calls, one per line. |
Returns
The circle.shot elements reconstructed from every matching line (non-matching lines, e.g. the addShot function definition line itself, are silently skipped).
Example
from sportsdataverse.mbb.mbb_ncaa_shot_parser import shot_js_to_html
js = "addShot(27.0, 77.0, 392, false, 1, 'title text', 'class', false);"
circles = shot_js_to_html(js)
start_time_from_period
start_time_from_period(period: 'int', is_women_game: 'bool') -> 'float'
The game-clock time (minutes elapsed) a period starts at
(ExtractorUtils.scala:272-281).
Women's games play four 10-minute quarters then 5-minute overtimes; men's games play two 20-minute halves then 5-minute overtimes.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
period | int | The 1-indexed period number (1/2 = halves for men, 1-4 = quarters for women, 5+ = overtimes for both). | |
is_women_game | bool | Whether to use the women's (quarters) or men's (halves) period schedule. |
Returns
The game-clock minute the period begins at.
Example
from sportsdataverse.mbb.mbb_ncaa_stints import start_time_from_period
start_time_from_period(2, is_women_game=False) # 20.0 (men's 2nd half)
start_time_from_period(1, is_women_game=True) # 0.0 (women's 1st quarter)
start_time_from_period(6, is_women_game=False) # 45.0 (men's 2nd OT)
sum_event_stats
sum_event_stats(lhs: 'LineupEventStats', rhs: 'LineupEventStats') -> 'LineupEventStats'
Field-wise add two :class:`~sportsdataverse.mbb.mbb_ncaa_models
.LineupEventStats (protected def sum_event_stats, LineupUtils.scala
:1534-1622, debug-only -- the Scala's own docstring says "just used for debug"). The Scala builds this via shapeless.Generic` field-zipping;
this port is an explicit field-by-field call since Python has no
equivalent generic-programming machinery.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
lhs | LineupEventStats | The left-hand stat tree. | |
rhs | LineupEventStats | The right-hand stat tree. |
Returns
A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every field summed (see the module's private sum_*helpers for theOptional`/nested-field summing rules).
Example
from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_event_stats
from sportsdataverse.mbb.mbb_ncaa_models import LineupEventStats
sum_event_stats(LineupEventStats.empty(), LineupEventStats.empty()).num_events
sum_shot_infos
sum_shot_infos(shot_infos: 'list[PlayerShotInfo]') -> 'Optional[PlayerShotInfo]'
Field-wise sum a list of :class:`~sportsdataverse.mbb.mbb_ncaa_models
.PlayerShotInfo\ s (sum_shot_infos, LineupUtils.scala:1625-1655`,
debug-only).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
shot_infos | list[PlayerShotInfo] | The list to combine, in order. |
Returns
None if shot_infos is empty; the single element if there's exactly one; otherwise a left-fold of pairwise field-wise sums (reduceOption).
Example
from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_shot_infos
from sportsdataverse.mbb.mbb_ncaa_models import PlayerShotInfo
sum_shot_infos([PlayerShotInfo(ast_3pm=(1, 0, 0, 0, 0)), PlayerShotInfo(ast_3pm=(0, 1, 0, 0, 0))])
td_at
td_at(row: 'Tag', n: 'int') -> 'Optional[Tag]'
JSoup row >?> element("td:eq(n)"): the n-th <td> child.
Soupsieve has no :eq() positional pseudo-class (unlike JSoup), so
this is a plain 0-indexed lookup into row.find_all("td"), guarded
against an out-of-range index (JSoup's >?> returns None rather
than raising when the selector matches nothing).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
row | Tag | The row (or other container) element to search. | |
n | int | The 0-indexed <td> position. |
Returns
The n-th <td> descendant, or None if row has fewer than n + 1 of them.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, td_at
soup = parse_html("<tr><td>A</td><td>B</td></tr>")
row = soup.find("tr")
td_at(row, 1).get_text() # "B"
td_at(row, 5) # None
transform_shot_location
transform_shot_location(x: 'float', y: 'float', second_half_switch: 'bool', team_shooting_left_in_first_period: 'bool', is_offensive: 'bool') -> 'tuple[float, float, float, float]'
Transforms a raw SVG pixel location into feet from the basket, always
oriented as if shooting towards the left goal (ShotEventParser .transform_shot_location, :588-620).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
x | float | Raw SVG cx pixel coordinate. | |
y | float | Raw SVG cy pixel coordinate. | |
second_half_switch | bool | Whether this shot is in the "other" half of the game from team_shooting_left_in_first_period (each False factor below flips which side is treated as "left"). | |
team_shooting_left_in_first_period | bool | Whether the team under analysis shot towards the left goal in the first period (see is_team_shooting_left_to_start). | |
is_offensive | bool | Whether the team under analysis is shooting (an opponent shot flips the expected side again). |
Returns
(x, y, alt_x, alt_y) in feet -- the believed-correct location, then the alternative (mirror-image) location, both relative to the goal the shot is (believed to be) attacking.
Example
from sportsdataverse.mbb.mbb_ncaa_shot_parser import transform_shot_location
transform_shot_location(310.2, 235, False, False, True)
update_config
update_config(**kwargs: 'object') -> 'NcaaFetchConfig'
Update the active config in place.
Returns
The (mutated) global config object.
Example
from sportsdataverse.mbb.mbb_ncaa_fetch import update_config
update_config(proxy_url="http://user:pass@1.2.3.4:8080")
validate_box_score
validate_box_score(team: 'TeamId', lineup: 'list[str]') -> 'Union[list[PlayerCodeId], ParseError]'
Checks there are no duplicates in the lineup (``BoxscoreParser
.validate_box_score, :388-404``).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
team | TeamId | The team the lineup belongs to (feeds ~sportsdataverse.mbb.mbb_ncaa_stints.build_player_code's team-scoped misspelling corrections). | |
lineup | list[str] | The raw player-name strings, in whatever order they were assembled by inject_validated_players. |
Returns
lineup, mapped to ~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (same order, no sort -- see the module docstring's "not sorted" note). When two teammates collide on the {first-two-letters}{Surname} scheme -- siblings, in practice -- only the colliding players are re-coded to {First}{Last} by disambiguate_sibling_codes; every other player keeps the Scala-faithful code. This is a DELIBERATE divergence from ExtractorUtils.scala, which rejects the game: since a team's roster is the same all season, one sibling pair cost the team its ENTIRE season of lineups. A ~sportsdataverse.mbb.mbb_ncaa_data_quality.ParseErroris returned only when widening cannot separate them, i.e. two players with the SAME full name -- genuinely ambiguous, so still an error. Callers must not re-derive a code from a name after this point:build_player_codewould undo the widening and silently drop one twin. Use~sportsdataverse.mbb.mbb_ncaa_names.code_from_box`, which resolves against this roster.
Example
from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import validate_box_score
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
validate_box_score(TeamId("Team"), ["Player One", "Player Two"])
validate_lineup
validate_lineup(lineup_event: 'LineupEvent', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'list[ValidationError]'
Flags a lineup stint as internally inconsistent, via 3 independent
checks (LineupErrorAnalysisUtils.validate_lineup, :181-218).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
lineup_event | LineupEvent | The lineup stint to validate. | |
box_lineup | LineupEvent | The team's box-score lineup event (players is the full roster) -- used both to build the name-resolution context (see ~sportsdataverse.mbb.mbb_ncaa_names.build_tidy_player_context) and, indirectly, as the source of players_out for jersey-number resolution inside ~sportsdataverse.mbb .mbb_ncaa_names.tidy_player. | |
valid_player_codes | set[str] | Every player code that's actually on the box score / roster for this team-season. |
Returns
The failing ValidationError\ s, in declaration order (see the module docstring's "Return shape" note) -- empty if lineup_event is clean. * ValidationError.WRONG_NUMBER_OF_PLAYERS -- lineup_event doesn't have exactly 5 players on the floor. * ValidationError.UNKNOWN_PLAYERS -- some player on the floor isn't in valid_player_codes. * ValidationError.INACTIVE_PLAYERS -- some player mentioned in lineup_event's own (team-side) raw game events resolves to a code not in valid_player_codes (i.e. isn't on the floor, per the lineup being validated).
Example
from sportsdataverse.mbb.mbb_ncaa_stint_validation import validate_lineup
errors = validate_lineup(lineup_event, box_lineup, {"MiMitchell", "BbBob"})
assert not errors # a clean lineup returns []