Skip to main content

MBB — additional Python functions — stats.ncaa.org: ncaa_mbb–validate_lineup

ncaa_mbb_shot_locations​

ncaa_mbb_shot_locations(game_ids: "'Sequence[object]'", *, fetcher: 'Optional[_SupportsFetchGameBox]' = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, Any]'"

Scrape MBB shot locations for one or more games (bigballR

get_shot_locations, get_shot_locations.R:3-89).

Fetches each game's stats.ncaa.org/contests/{id}/box_score page and parses the embedded shot-chart JS through parse_ncaa_bb_shots. NA ids are dropped up front (R :5); per-game "shots found" messages go to the module logger (R message, :69-70).

Parameters

ParameterTypeDefaultDescription
game_idsSequence[object]NCAA contest ids; None/NaN entries are dropped.
fetcherOptional[_SupportsFetchGameBox]NoneOptional injected fetcher exposing fetch_game_box (for tests/offline use). Defaults to a fresh NcaaFetcher.with_browser() context per call.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

All games' shots row-bound (zero-row SHOTS_SCHEMA frame when no ids survive or no charts are found).

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
clockcharacterGame clock value.
game_secondsinteger
teamcharacterTeam-side label or team identifier.
playercharacterPlayer name.
shot_resultcharacterShot result ('Made' / 'Missed').
xdoubleX.
ydoubleY.
shot_distdouble

Example

from sportsdataverse.mbb.mbb_ncaa_shots import ncaa_mbb_shot_locations
df = ncaa_mbb_shot_locations(["6470186", "6479639"])
print(df.shape)

# Offline with an injected fetcher

df = ncaa_mbb_shot_locations(["6470186"], fetcher=my_fetcher)

# Pipeline next step (one line)

df.group_by("team").agg(pl.col("shot_dist").mean()).head()

ncaa_mbb_team_stats​

ncaa_mbb_team_stats(pbp: 'pl.DataFrame', *, include_transition: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into per-team game stats.

Port of bigballR get_team_stats (all_functions.R:2530-2538): the ten on-court columns are blanked so every row shares one "lineup", then get_lineups (ncaa_mbb_lineups) runs per game and the lineup key columns are dropped — yielding two rows (one per team) per game.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract. May span multiple games.
include_transitionboolFalseWhen True, append the trans/half split surface plus o_trans_pct/d_trans_pct.
fix_tip_inboolTrueWhen True (default), rim stats count the scrape engine's real "Tip In" vocabulary; False reproduces R's literal "Tip-In" bug for oracle parity.
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with one row per team per game — TEAM_STATS_COLUMNS (73) or TEAM_STATS_TRANSITION_COLUMNS with include_transition=True. Games ordered by the Utf8 game_id byte sort (R's do() sorts a numeric ID — identical for equal-width ids), teams within a game byte-sorted. Empty input yields an empty frame with the documented schema.

col_nametypedescription
game_idcharacterUnique game identifier.
homecharacterHome.
awaycharacterAway record.
teamcharacterTeam-side label or team identifier.
minsdouble
o_minsdouble
d_minsdouble
o_possdouble
d_possdouble
ortgdouble
drtgdouble
netrtgdouble
ptsdoublePoints scored.
d_ptsdouble
fgadoubleField goal attempts.
d_fgadouble
fgmdoubleField goals made.
d_fgmdouble
tpadouble
d_tpadouble
tpmdouble
d_tpmdouble
ftadoubleFree throw attempts.
d_ftadouble
ftmdoubleFree throws made.
d_ftmdouble
rimadouble
d_rimadouble
rimmdouble
d_rimmdouble
orbdouble
d_orbdouble
drbdouble
d_drbdouble
blkdoubleBlocks.
d_blkdouble
todoubleTo.
d_todouble
astdoubleAssists.
d_astdouble
e_possdouble
fg_pctdoubleField goal percentage (0-1).
d_fg_pctdouble
tppdouble
d_tppdouble
ftpdouble
d_ftpdouble
efg_pctdouble
d_efg_pctdouble
ts_pctdoubleTrue shooting percentage (0-1).
d_ts_pctdouble
rim_pctdouble
d_rim_pctdouble
mid_pctdouble
d_mid_pctdouble
tp_ratedouble
d_tp_ratedouble
rim_ratedouble
d_rim_ratedouble
mid_ratedouble
d_mid_ratedouble
ft_ratedoubleFt rate.
d_ft_ratedouble
ast_ratedouble
d_ast_ratedouble
to_ratedoubleTo rate.
d_to_ratedouble
blk_ratedouble
o_blk_ratedouble
orb_pctdoubleOffensive rebound percentage.
drb_pctdoubleDefensive rebound percentage.
time_per_possdouble
d_time_per_possdouble

Example

from sportsdataverse.mbb.mbb_ncaa_stats_agg import ncaa_mbb_team_stats
teams = ncaa_mbb_team_stats(pbp)
print(teams.shape)

# Transition splits, pandas out

df_pd = ncaa_mbb_team_stats(pbp, include_transition=True, return_as_pandas=True)

# Pipeline next step (one line)

teams.sort("netrtg", descending=True).head()

phase1_shot_event_enrichment​

phase1_shot_event_enrichment(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]', second_half_override: 'Optional[set[int]]' = None) -> 'list[ShotEvent]'

The court-geometry enrichment pass: ascending time, coordinate

transform + geo synthesis, and the self-correcting side-flip re-run (ShotEventParser.phase1_shot_event_enrichment, :415-528).

For each shot: compute the ascending game time, decide (from is_team_shooting_left_to_start + which half the period falls in) whether the shot's side needs flipping, run transform_shot_location to get both the believed-correct and alternative (mirrored) locations, keep whichever is closer to the basket (a >1.2x distance advantage for the "alternative" wins, or ANY shot taken with <0.1 min left on the clock always keeps the original -- a half-court heave near the buzzer is plausible, so the tie-break favors trusting the raw geometry there), then synthesize a lat/lon.

After all shots are processed, if any period had >=6 shots AND more than 75% of them came back implausibly long-distance (>50ft), the whole pass re-runs ONCE with those periods' orientation flipped (the self-correcting part) -- second_half_override is None on the initial call and a non-None set on the one allowed retry, preventing infinite recursion.

Parameters

ParameterTypeDefaultDescription
sorted_very_raw_eventslist[tuple[int, ShotEvent]]The chronologically-sorted (period, shot) pairs from parse_shot_html, pre-geometry-transform.
second_half_overrideOptional[set[int]]NoneThe set of periods whose second_half_switch orientation should be inverted (the self-correction re-run's input); None on the first call.

Returns

The fully court-geometry-enriched shots, in the same order as sorted_very_raw_events.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import phase1_shot_event_enrichment
shots = phase1_shot_event_enrichment([(1, very_raw_shot)])

playwright_transport​

playwright_transport(*, headless_new: 'bool' = True, challenge_wait_ms: 'int' = 8000, nav_timeout_ms: 'int' = 45000, user_agent: 'Optional[str]' = None, solve_attempts: 'int' = 3, relaunch_backoff: 'float' = 2.0) -> "'_PlaywrightTransport'"

Build the suggested stats.ncaa.org game-detail scraping transport.

Drives a real Chromium via Playwright in Chrome's new-headless mode (--headless=new) to clear the Akamai bm-verify challenge that curl_cffi cannot, then serves raw server HTML for the 5a-5e parsers. Playwright is a lazy optional import (not a hard dependency); a clear ImportError fires on first use if it is missing.

Parameters

ParameterTypeDefaultDescription
headless_newboolTrueUse --headless=new (real-GPU render, no window) -- the default and the proven-working mode. False runs old headless (headless_shell), which Akamai flags -- avoid.
challenge_wait_msint8000Milliseconds to let the bm-verify sensor run after the first navigation.
nav_timeout_msint45000Per-navigation timeout.
user_agentOptional[str]NoneOverride the Chrome UA string.
solve_attemptsint3
relaunch_backofffloat2.0

Returns

A stateful, callable FetchTransport reusing one browser for the session. Close it when done (it is a context manager, has close(), and registers an atexit safety net).

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import NcaaFetcher
with NcaaFetcher.with_browser() as fetcher:
pbp = fetcher.fetch_game_pbp("1613299") # raw PBP HTML
box = fetcher.fetch_game_individual_stats("1613299") # raw box HTML
# -> feed to get_box_lineup / create_lineup_data (mbb_ncaa_*_parser)

remove_diacritics​

remove_diacritics(fragment: 'str') -> 'str'

Strip diacritical marks, e.g. "Juhász" -> "Juhasz"

(ExtractorUtils.scala:38-43: NFD normalization then removal of the combining-diacritical-marks block).

Parameters

ParameterTypeDefaultDescription
fragmentstrAny string (a full player name or a name fragment).

Returns

The string with combining marks removed.

Example

from sportsdataverse.mbb.mbb_ncaa_stints import remove_diacritics
print(remove_diacritics("Dorka Juhász")) # "Dorka Juhasz"

remove_html_encoding​

remove_html_encoding(html_str: 'str') -> 'str'

Undo a handful of literal HTML entity escapes (``ExtractorUtils

.remove_html_encoding, ExtractorUtils.scala:25-33). **Scope addition, Task 5e.5** -- the first consumer is mbb_ncaa_shot_parser.parse_shot_html(theplayername / shooting team name extracted from an SVG shot's` text).

In practice bs4/lxml already decode standard HTML entities (&#39;, &quot;, `&``) while parsing text nodes, so this is usually a no-op by the time it runs on already-parsed text -- ported anyway for exact behavioral parity with any double-escaped input the upstream Scala guards against (JSoup has the same auto-decoding behavior, so the Scala original is equally a defensive no-op in the common case).

Parameters

ParameterTypeDefaultDescription
html_strstrAny string, typically already-parsed element text.

Returns

html_str with &#39;/&quot;/&amp; replaced by their literal characters, only if "&" appears at all (short-circuit matching the Scala's if (html_str.indexOf("&") >= 0) guard).

Example

from sportsdataverse.mbb.mbb_ncaa_stints import remove_html_encoding
remove_html_encoding("De&#39;Shayne") # "De'Shayne"
remove_html_encoding("Plain Name") # "Plain Name" (unchanged)

reorder_and_reverse​

reorder_and_reverse(reversed_partial_events: 'Iterable[PlayByPlayEvent]') -> 'list[PlayByPlayEvent]'

Orders same-minute play-by-play events so subs never enclose the plays

they logically precede/follow (ExtractorUtils.scala:435-599).

Groups consecutive events sharing the same min into a block (the input arrives in descending/reverse-chronological order, so blocks are discovered and internally accumulated in reverse too), then -- for any block containing a sub -- reorders it via inner_sort: events referencing a subbed-OUT player (or scoring no higher than the sub) land in a pre-sub group, the subs themselves come next (in ascending-score order), and events referencing a subbed-IN player (or scoring higher than the sub) land in a trailing post-sub group. Free-throw attempts sharing the sub's inferred "direction" (team vs. opponent, inferred from the nearest preceding shot/FT/foul) are pulled into the pre-sub group unless the shooter is one of the players being subbed in. Blocks with no sub are returned unchanged apart from the initial score-based sort.

Parameters

ParameterTypeDefaultDescription
reversed_partial_eventsIterable[PlayByPlayEvent]Events for one lineup event, in reverse-chronological (descending-time) order -- the natural order encountered walking play-by-play text bottom-up.

Returns

The same events, forward-chronological (ascending time), with each same-minute block internally reordered so no sub encloses a play it logically shouldn't.

Example

from sportsdataverse.mbb.mbb_ncaa_models import Score
from sportsdataverse.mbb.mbb_ncaa_stints import (
OtherTeamEvent,
SubInEvent,
reorder_and_reverse,
)
events = [
SubInEvent(0.4, Score(0, 0), "player1"),
OtherTeamEvent(0.4, Score(0, 0), "rebound"),
]
reorder_and_reverse(events)
# [OtherTeamEvent(...), SubInEvent(...)]

reset_config​

reset_config() -> 'NcaaFetchConfig'

Reset the active config to its env-var-derived defaults.

Returns

The live singleton, now holding the env-var-derived defaults again.

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import update_config, reset_config
update_config(timeout=5)
reset_config()

right_kind_of_shot​

right_kind_of_shot(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', strict: 'bool') -> 'bool'

Whether pbp_event's shot type is compatible with shot's

distance and make/miss (ShotEnrichmentUtils.right_kind_of_shot, PlayByPlayUtils.scala:659-679).

The distance-in-the-data is approximate, so exact 2-vs-3 discrimination is impossible; this only rules out the obvious mismatches (a clearly-short shot matched to a 3, or vice versa) and always requires make/miss agreement.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched (pts/dist read).
pbp_eventMiscGameEventThe candidate play-by-play event.
strictboolIf True, also apply the distance gate; if False, only the make/miss agreement is required.

Returns

True if the event could plausibly be this shot.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import right_kind_of_shot
right_kind_of_shot(shot, pbp_event, strict=True)

run_iterative_adjustment_with_hca​

run_iterative_adjustment_with_hca(teams: 'Sequence[TeamDetail]', team_by_name: 'dict[str, TeamDetail]', fields: 'Sequence[str]', league_averages: 'LeagueAverages', poss_splits: 'dict[str, PossessionSplits]', *, max_iterations: 'int' = 100, tolerance: 'float' = 1e-06) -> 'IterationResult'

KenPom-style SoS + HCA fixed-point solver (runIterativeAdjustmentWithHCA, ts:306-527).

Each iteration (Jacobi -- all teams read the previous iteration's adjustments, then commit together):

  1. Per team/field, adjust every game adj_game = raw_game * (league / (opp_adj +/- hca)) and take the weighted mean; a field with no valid games keeps its current value.
  2. Re-estimate per-field HCA from home/away possession-imbalance residuals hca = sum((raw - pred) * |imbalance|) / sum(|imbalance|) over teams with |imbalance| >= IMBALANCE_MIN.

Stops when the max per-team/field change drops below tolerance or after max_iterations sweeps (the HCA re-estimate still runs on the final sweep). The cross-guard on the per-game branch, the asymmetric residual prediction, and the cross-named opponent strengths are all preserved -- see the module docstring's landmine list.

Parameters

ParameterTypeDefaultDescription
teamsSequence[TeamDetail]The teams to solve over.
team_by_namedict[str, TeamDetail]{team_name: team_detail} for opponent lookup.
fieldsSequence[str]The stat fields to solve.
league_averagesLeagueAveragesOutput of compute_league_averages_from_per_game.
poss_splitsdict[str, PossessionSplits]{team_name: PossessionSplits }.
max_iterationsint100Iteration cap (default MAX_ITERATIONS; pin to 1 to inspect a single sweep).
tolerancefloat1e-06Convergence tolerance (default TOLERANCE).

Returns

An IterationResult (adj_values, hca_per_field).

Example

from sportsdataverse.mbb.mbb_ncaa_strength import (
STRENGTH_ADJUSTED_FIELDS,
compute_league_averages_from_per_game,
compute_possession_splits,
run_iterative_adjustment_with_hca,
)

by_name = {t["team_name"]: t for t in teams}
league = compute_league_averages_from_per_game(teams)
splits = {t["team_name"]: compute_possession_splits(t) for t in teams}
result = run_iterative_adjustment_with_hca(
teams, by_name, STRENGTH_ADJUSTED_FIELDS, league, splits,
)
print(result.hca_per_field["3p"]["hca_off"])

same_school​

same_school(a: 'str', b: 'str') -> 'bool'

Whether two team-name spellings denote the same school.

Exact match, or both names inside one team_name_equivalents class.

Parameters

ParameterTypeDefaultDescription
astrOne team-name spelling.
bstrThe other team-name spelling.

Returns

True if the two names refer to the same school.

Example

from sportsdataverse.mbb.mbb_ncaa_data_quality import same_school
same_school("New Orleans", "LSU New Orleans") # True
same_school("Miami (FL)", "Miami (OH)") # False

select_contains​

select_contains(root: 'Tag', selector: 'str', text: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":contains(text)"): candidates whose full

text (own + every descendant's) case-insensitively CONTAINS text as a plain substring -- not a regex (Task 5e.2 addition; see the module docstring's "Critical divergence" note).

JSoup's :contains() is documented case-insensitive substring containment; soupsieve's :-soup-contains() (the non-deprecated spelling of its :contains()) is case-SENSITIVE, with no case-insensitive variant of its own. Reproducing JSoup's actual semantics therefore needs this helper rather than :-soup-contains().

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector for the structural part of the match (everything before :contains).
textstrThe plain substring each candidate's collapsed text must case-insensitively contain.

Returns

Every selector match whose jsoup_text case-insensitively contains text, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_contains
soup = parse_html("<td>game date:</td><td>Location:</td>")
select_contains(soup, "td", "Game Date:") # [<td>game date:</td>]

select_matching​

select_matching(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":matches(regex)"): candidates whose full

text (own + every descendant's) matches regex.

Soupsieve has no :matches() pseudo-class equivalent, so this runs the plain structural selector first, then filters by re.search over each candidate's jsoup_text (own text plus descendants', matching JSoup's :matches() semantics -- as opposed to select_matching_own's own-text-only :matchesOwn()).

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector.
regexstrThe pattern each candidate's collapsed text must re.search-match.

Returns

Every selector match whose jsoup_text contains a regex match, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching
soup = parse_html("<div><p>Home Team</p><p>Away Team</p></div>")
select_matching(soup, "p", r"^Home") # [<p>Home Team</p>]

select_matching_own​

select_matching_own(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":matchesOwn(regex)"): candidates whose

OWN text only (excluding descendant elements' text) matches regex.

JSoup's Element.ownText() walks only the element's direct TextNode children, not text nested inside child elements -- the same distinction bs4 draws between a tag's direct bs4.NavigableString children and its full .get_text().

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector.
regexstrThe pattern each candidate's own (whitespace-collapsed) text must re.search-match.

Returns

Every selector match whose own text contains a regex match, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching_own
soup = parse_html('<div class="card-header">Coach <b>Info</b></div>')
select_matching_own(soup, "div.card-header", r"^Coach")
# [<div class="card-header">Coach <b>Info</b></div>]

shot_js_to_html​

shot_js_to_html(js: 'str') -> 'list[Tag]'

Converts client-side addShot(...) JS calls into parseable

circle.shot HTML, for pages where the shot map is built on the fly rather than baked into the initial HTML (ShotEventParser .shot_js_to_html, :266-283). See the module docstring's "Scala idiom decision" note -- the Scala's builders/browser parameters are dropped here since the Scala body never actually uses them.

Parameters

ParameterTypeDefaultDescription
jsstrThe concatenated <script> text containing one or more addShot(x, y, ..., 'title', ...) calls, one per line.

Returns

The circle.shot elements reconstructed from every matching line (non-matching lines, e.g. the addShot function definition line itself, are silently skipped).

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import shot_js_to_html
js = "addShot(27.0, 77.0, 392, false, 1, 'title text', 'class', false);"
circles = shot_js_to_html(js)

start_time_from_period​

start_time_from_period(period: 'int', is_women_game: 'bool') -> 'float'

The game-clock time (minutes elapsed) a period starts at

(ExtractorUtils.scala:272-281).

Women's games play four 10-minute quarters then 5-minute overtimes; men's games play two 20-minute halves then 5-minute overtimes.

Parameters

ParameterTypeDefaultDescription
periodintThe 1-indexed period number (1/2 = halves for men, 1-4 = quarters for women, 5+ = overtimes for both).
is_women_gameboolWhether to use the women's (quarters) or men's (halves) period schedule.

Returns

The game-clock minute the period begins at.

Example

from sportsdataverse.mbb.mbb_ncaa_stints import start_time_from_period
start_time_from_period(2, is_women_game=False) # 20.0 (men's 2nd half)
start_time_from_period(1, is_women_game=True) # 0.0 (women's 1st quarter)
start_time_from_period(6, is_women_game=False) # 45.0 (men's 2nd OT)

sum_event_stats​

sum_event_stats(lhs: 'LineupEventStats', rhs: 'LineupEventStats') -> 'LineupEventStats'

Field-wise add two :class:`~sportsdataverse.mbb.mbb_ncaa_models

.LineupEventStats (protected def sum_event_stats, LineupUtils.scala :1534-1622, debug-only -- the Scala's own docstring says "just used for debug"). The Scala builds this via shapeless.Generic` field-zipping; this port is an explicit field-by-field call since Python has no equivalent generic-programming machinery.

Parameters

ParameterTypeDefaultDescription
lhsLineupEventStatsThe left-hand stat tree.
rhsLineupEventStatsThe right-hand stat tree.

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every field summed (see the module's private sum_*helpers for theOptional`/nested-field summing rules).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_event_stats
from sportsdataverse.mbb.mbb_ncaa_models import LineupEventStats

sum_event_stats(LineupEventStats.empty(), LineupEventStats.empty()).num_events

sum_shot_infos​

sum_shot_infos(shot_infos: 'list[PlayerShotInfo]') -> 'Optional[PlayerShotInfo]'

Field-wise sum a list of :class:`~sportsdataverse.mbb.mbb_ncaa_models

.PlayerShotInfo\ s (sum_shot_infos, LineupUtils.scala:1625-1655`, debug-only).

Parameters

ParameterTypeDefaultDescription
shot_infoslist[PlayerShotInfo]The list to combine, in order.

Returns

None if shot_infos is empty; the single element if there's exactly one; otherwise a left-fold of pairwise field-wise sums (reduceOption).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_shot_infos
from sportsdataverse.mbb.mbb_ncaa_models import PlayerShotInfo

sum_shot_infos([PlayerShotInfo(ast_3pm=(1, 0, 0, 0, 0)), PlayerShotInfo(ast_3pm=(0, 1, 0, 0, 0))])

td_at​

td_at(row: 'Tag', n: 'int') -> 'Optional[Tag]'

JSoup row >?> element("td:eq(n)"): the n-th <td> child.

Soupsieve has no :eq() positional pseudo-class (unlike JSoup), so this is a plain 0-indexed lookup into row.find_all("td"), guarded against an out-of-range index (JSoup's >?> returns None rather than raising when the selector matches nothing).

Parameters

ParameterTypeDefaultDescription
rowTagThe row (or other container) element to search.
nintThe 0-indexed <td> position.

Returns

The n-th <td> descendant, or None if row has fewer than n + 1 of them.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, td_at
soup = parse_html("<tr><td>A</td><td>B</td></tr>")
row = soup.find("tr")
td_at(row, 1).get_text() # "B"
td_at(row, 5) # None

transform_shot_location​

transform_shot_location(x: 'float', y: 'float', second_half_switch: 'bool', team_shooting_left_in_first_period: 'bool', is_offensive: 'bool') -> 'tuple[float, float, float, float]'

Transforms a raw SVG pixel location into feet from the basket, always

oriented as if shooting towards the left goal (ShotEventParser .transform_shot_location, :588-620).

Parameters

ParameterTypeDefaultDescription
xfloatRaw SVG cx pixel coordinate.
yfloatRaw SVG cy pixel coordinate.
second_half_switchboolWhether this shot is in the "other" half of the game from team_shooting_left_in_first_period (each False factor below flips which side is treated as "left").
team_shooting_left_in_first_periodboolWhether the team under analysis shot towards the left goal in the first period (see is_team_shooting_left_to_start).
is_offensiveboolWhether the team under analysis is shooting (an opponent shot flips the expected side again).

Returns

(x, y, alt_x, alt_y) in feet -- the believed-correct location, then the alternative (mirror-image) location, both relative to the goal the shot is (believed to be) attacking.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import transform_shot_location
transform_shot_location(310.2, 235, False, False, True)

update_config​

update_config(**kwargs: 'object') -> 'NcaaFetchConfig'

Update the active config in place.

Returns

The (mutated) global config object.

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import update_config
update_config(proxy_url="http://user:pass@1.2.3.4:8080")

validate_box_score​

validate_box_score(team: 'TeamId', lineup: 'list[str]') -> 'Union[list[PlayerCodeId], ParseError]'

Checks there are no duplicates in the lineup (``BoxscoreParser

.validate_box_score, :388-404``).

Parameters

ParameterTypeDefaultDescription
teamTeamIdThe team the lineup belongs to (feeds ~sportsdataverse.mbb.mbb_ncaa_stints.build_player_code's team-scoped misspelling corrections).
lineuplist[str]The raw player-name strings, in whatever order they were assembled by inject_validated_players.

Returns

lineup, mapped to ~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (same order, no sort -- see the module docstring's "not sorted" note). When two teammates collide on the {first-two-letters}{Surname} scheme -- siblings, in practice -- only the colliding players are re-coded to {First}{Last} by disambiguate_sibling_codes; every other player keeps the Scala-faithful code. This is a DELIBERATE divergence from ExtractorUtils.scala, which rejects the game: since a team's roster is the same all season, one sibling pair cost the team its ENTIRE season of lineups. A ~sportsdataverse.mbb.mbb_ncaa_data_quality.ParseErroris returned only when widening cannot separate them, i.e. two players with the SAME full name -- genuinely ambiguous, so still an error. Callers must not re-derive a code from a name after this point:build_player_codewould undo the widening and silently drop one twin. Use~sportsdataverse.mbb.mbb_ncaa_names.code_from_box`, which resolves against this roster.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import validate_box_score
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
validate_box_score(TeamId("Team"), ["Player One", "Player Two"])

validate_lineup​

validate_lineup(lineup_event: 'LineupEvent', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'list[ValidationError]'

Flags a lineup stint as internally inconsistent, via 3 independent

checks (LineupErrorAnalysisUtils.validate_lineup, :181-218).

Parameters

ParameterTypeDefaultDescription
lineup_eventLineupEventThe lineup stint to validate.
box_lineupLineupEventThe team's box-score lineup event (players is the full roster) -- used both to build the name-resolution context (see ~sportsdataverse.mbb.mbb_ncaa_names.build_tidy_player_context) and, indirectly, as the source of players_out for jersey-number resolution inside ~sportsdataverse.mbb .mbb_ncaa_names.tidy_player.
valid_player_codesset[str]Every player code that's actually on the box score / roster for this team-season.

Returns

The failing ValidationError\ s, in declaration order (see the module docstring's "Return shape" note) -- empty if lineup_event is clean. * ValidationError.WRONG_NUMBER_OF_PLAYERS -- lineup_event doesn't have exactly 5 players on the floor. * ValidationError.UNKNOWN_PLAYERS -- some player on the floor isn't in valid_player_codes. * ValidationError.INACTIVE_PLAYERS -- some player mentioned in lineup_event's own (team-side) raw game events resolves to a code not in valid_player_codes (i.e. isn't on the floor, per the lineup being validated).

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import validate_lineup
errors = validate_lineup(lineup_event, box_lineup, {"MiMitchell", "BbBob"})
assert not errors # a clean lineup returns []