Skip to main content

MBB — additional Python functions — Calc

calc_collinearity_diag​

calc_collinearity_diag(weight_matrix: 'NDArray[np.float64]', ctx: 'RapmPlayerContext') -> 'RapmPreProcDiagnostics'

Multi-collinearity diagnostic between the players in an off/def design matrix.

Faithful port of RapmUtils.calcCollinearityDiag (RapmUtils.ts:1629-1760). Runs an SVD of weight_matrix, builds condition indices ("lineup combos") from the ratio of the largest to each singular value, and a variance-decomposition-proportions ("VDP") matrix identifying which players load onto which collinear combo -- the classic Belsley-Kuh-Welsch collinearity-diagnostics recipe (see the upstream comment's colldiag.m citation). Also builds a plain Pearson player/player correlation matrix (calc_player_correlations) and folds it into a possession -weighted adaptive_correl_weights` summary per player.

numpy.linalg.svd(weight_matrix, full_matrices=False) replaces svd-js's SVD(weightMatrix, false). Both are the standard Golub-Kahan-Reinsch decomposition (A = U @ diag(S) @ Vᵀ); numpy's Vh return value already is Vᵀ (what the TS code separately computes via transpose(matrix(v))), so this port skips that transpose. The TS code (and this port) never reads u/the first SVD return -- only q/S (singular values) and v/Vᵀ. Singular -vector sign is immaterial here: every place V is used (phiMatrix/phi_matrix) squares each entry (val * val), and a per-singular-value sign flip on U/V together is a valid SVD regardless -- so any U/V sign convention difference between svd-js and LAPACK (numpy's backend) cannot change this function's output. Singular-value ordering is likewise immaterial: both this port and the TS source explicitly re-sort q (ascending, carrying the original index along) before using it, so whichever order either SVD implementation returns values in, the final result only depends on the values themselves (up to the explicit resort), not on numpy's native descending convention vs whatever order svd-js happens to return.

correl_matrix/poss_correl_matrix stay numpy.ndarray (see the module docstring's "Task 3.6 notes" for why this doesn't hit the Task 3.5 "ndarray breaks deep ==" concern).

Parameters

ParameterTypeDefaultDescription
weight_matrixNDArray[float64]An off/def design matrix, shape (num_lineups, ctx["num_players"]) (e.g. calc_player_weights's first return value, or a hand-built matrix for isolated testing).
ctxRapmPlayerContextA RapmPlayerContext. ctx["num_players"] sizes every per-player structure; ctx["col_to_player"] keys player_combos.

Returns

A RapmPreProcDiagnostics.

Example

from sportsdataverse.mbb.mbb_rapm import calc_collinearity_diag, calc_player_weights

off_weights, _ = calc_player_weights(ctx)
diag = calc_collinearity_diag(off_weights, ctx)
print(diag["lineup_combos"][0]) # the worst-conditioned combo

calc_def_player_luck_adj​

calc_def_player_luck_adj(sample: 'LineupStatSet', base: 'LineupStatSet', avg_eff: 'float') -> 'DefLuckAdjustmentDiags'

Defensive 3P-luck adjustment for a single player.

Faithful port of LuckUtils.calcDefPlayerLuckAdj (LuckUtils.ts:402-426). Unlike calc_off_player_luck_adj, this is not a pure delegation -- see the module docstring's calc_def_player_luck_adj note for the translate() remap this wraps around calc_def_team_luck_adj.

Parameters

ParameterTypeDefaultDescription
sampleLineupStatSetThe player's stat dict for the period being luck-adjusted (must carry oppo_total_def_3p_made/oppo_total_def_3p_attempts -- there is no player-level def_3p field upstream, hence the remap).
baseLineupStatSetThe player's stat dict for the baseline/reference period.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).

Returns

Same shape as calc_def_team_luck_adj, computed against the translated (oppo_* -> def_*) player stat dicts.

Example

from sportsdataverse.mbb.mbb_luck import calc_def_player_luck_adj

diags = calc_def_player_luck_adj(sample_player, base_player, 100.0)
print(diags["deltaDefAdjEff"])

calc_def_team_luck_adj​

calc_def_team_luck_adj(sample: 'LineupStatSet', base: 'LineupStatSet', avg_eff: 'float', sample_def_3pa_override: 'float | None' = None) -> 'DefLuckAdjustmentDiags'

Defensive 3P-luck adjustment for a team (or lineup).

Faithful port of LuckUtils.calcDefTeamLuckAdj (LuckUtils.ts:429-531). See the module docstring for the SoS-vs-luck-split formula (LUCK_PCT) and the shared unguarded-division landmine.

Parameters

ParameterTypeDefaultDescription
sampleLineupStatSetThe team/lineup/player stat dict for the period being luck-adjusted (e.g. an on/off split or a single lineup).
baseLineupStatSetThe team/lineup/player stat dict for the baseline/reference period.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).
sample_def_3pa_overridefloat | NoneNoneWhen given, used as sampleDef3PA instead of sample["total_def_3p_attempts"] -- see calc_off_team_luck_adj's sample_3pa_override docstring for the shared "lineup regression" rationale (LuckUtils.ts:433-434, verbatim comment).

Returns

A DefLuckAdjustmentDiags dict -- TS-verbatim keys (avgEff, luckPct, baseDef3P, baseDef3PSos, baseDef3PA, basePoss, base3PSosAdj, sampleDef3P, sampleDef3PSos, sampleDef3PA, samplePoss, sample3PSosAdj, sampleDefEfg, sampleDefPpp, sampleOffSos, sampleDef3PRate, sampleDefFGA, sampleDefOrb, avg3PSosAdj, adjDef3P, delta3P, deltaDefEfg, deltaDefPppNoOrb, deltaMissesPct, deltaDefOrbFactor, deltaPtsOffMisses, deltaDefPpp, deltaDefAdjEff).

Example

from sportsdataverse.mbb.mbb_luck import calc_def_team_luck_adj

diags = calc_def_team_luck_adj(sample_team_off, base_team, 100.0)
print(diags["deltaDefAdjEff"])

calc_lineup_outputs​

calc_lineup_outputs(field: 'str', off_offset: 'float', def_offset: 'float', ctx: 'RapmPlayerContext', adaptive_correl_weights: 'list[float] | None' = None, use_old_val_if_possible: 'tuple[bool, bool]' = (False, False)) -> 'list[NDArray[np.float64]]'

Build the off/def target vectors the RAPM design matrices are fit against.

Faithful port of RapmUtils.calcLineupOutputs (RapmUtils.ts:598-751). For each filtered lineup, computes a possession-weighted residual: the lineup's own stat value, plus any global luck adjustment, minus the accumulated "prior offset" contributed by every player on the lineup (a strong-prior blend for kept players -- see get_strong_weight` -- or a fixed baseline contribution for removed players).

Upstream keeps this as a plain Array<Array<number>> (not a mathjs Matrix, unlike calc_player_weights's offWeights/ defWeights -- RapmUtils.test.ts's own tidyResults helper for this function has a visibly different shape, see the classification map in tests/fixtures/hoop_explorer/README.md). This port still materializes both output vectors as numpy.ndarray for consistency with calc_player_weights at the same dict -> array boundary -- Task 3.4's ridge-regression solve consumes both as arrays regardless of the upstream distinction.

Parameters

ParameterTypeDefaultDescription
fieldstrThe stat suffix to read off each lineup, e.g. "adj_ppp" (read as {prefix}_{field}, e.g. "off_adj_ppp").
off_offsetfloatThe D1-average offensive value for field (the regression's starting/baseline value on the RHS).
def_offsetfloatThe D1-average defensive value for field.
ctxRapmPlayerContextA RapmPlayerContext, e.g. from build_player_context.
adaptive_correl_weightslist[float] | NoneNoneOptional per-player adaptive-correlation weights (index-aligned with ctx["col_to_player"]), used as the strong-prior blend fallback when ctx["prior_info"] ["strong_weight"] < 0 -- see get_strong_weight`.
use_old_val_if_possibletuple[bool, bool](False, False)(use_old_val_for_off, use_old_val_for_def) -- whether to prefer each lineup/team stat's luck-adjusted old_value over its raw value when present. This is the luck-adjustment hook Task 3.1's classification map flags as an inherited coverage gap: the vendored oracle fixture has old_value == value on every field (via insertOldValues), so neither jest nor this port's replay test ever observes this flag change the resulting numbers -- only that passing it doesn't crash. See the module docstring's "Task 3.3 coverage gap" note.

Returns

[off_outputs, def_outputs] -- two 1-D numpy.ndarray target vectors, index-aligned with ctx["filtered_lineups"]("off"/"def") (plus one extra element each when ctx["unbias_weight"] > 0, an "unbiasing observation" target -- always unreached in production, same as calc_player_weights's extra row).

Example

from sportsdataverse.mbb.mbb_rapm import calc_lineup_outputs

off_outputs, def_outputs = calc_lineup_outputs(
"adj_ppp", 100.0, 100.0, ctx
)
print(off_outputs.shape) # (num_off_lineups,)

# Luck-adjusted variant (reads ``old_value`` where present)

off_luck, def_luck = calc_lineup_outputs(
"adj_ppp", 100.0, 100.0, ctx, use_old_val_if_possible=(True, True)
)

calc_off_player_luck_adj​

calc_off_player_luck_adj(sample_player: 'LineupStatSet', base_player: 'LineupStatSet', avg_eff: 'float') -> 'OffLuckAdjustmentDiags'

Offensive 3P-luck adjustment for a single player.

Faithful port of LuckUtils.calcOffPlayerLuckAdj (LuckUtils.ts:174-187). Per Task 2.1's surprise #4, this is a literal 1-player-team delegation to calc_off_team_luck_adj -- ORB effects are ignored for an individual player (the upstream comment: "the team calc basically works fine here, apart from ORBs, which we'll ignore").

Parameters

ParameterTypeDefaultDescription
sample_playerLineupStatSetThe player's stat dict for the period being luck-adjusted.
base_playerLineupStatSetThe player's stat dict for the baseline/reference period.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).

Returns

Same shape as calc_off_team_luck_adj -- identical to calling that function with sample_players=[sample_player], base_players_map={base_player["key"]: base_player}.

Example

from sportsdataverse.mbb.mbb_luck import calc_off_player_luck_adj

diags = calc_off_player_luck_adj(sample_player, base_player, 100.0)
print(diags["deltaOffAdjEff"])

calc_off_team_luck_adj​

calc_off_team_luck_adj(sample_team: 'LineupStatSet', sample_players: 'list[LineupStatSet]', base_team: 'LineupStatSet', base_players_map: 'dict[str, LineupStatSet]', avg_eff: 'float', sample_3pa_override: 'float | None' = None, manual_overrides: 'list[ManualOverride] | None' = None) -> 'OffLuckAdjustmentDiags'

Offensive 3P-luck adjustment for a team (or lineup).

Faithful port of LuckUtils.calcOffTeamLuckAdj (LuckUtils.ts:190-399). See the module docstring for the Bayesian-shrink formula, the JS-array- truthiness / object-selection landmines, and the one unguarded-division landmine this function carries.

Parameters

ParameterTypeDefaultDescription
sample_teamLineupStatSetThe team/lineup stat dict for the period being luck-adjusted (e.g. an on/off split or a single lineup).
sample_playerslist[LineupStatSet]The roster of per-player stat dicts backing sample_team (samplePlayers == players.map(on/off/baseline) per the upstream comment).
base_teamLineupStatSetThe team stat dict for the baseline/reference period (typically full-season).
base_players_mapdict[str, LineupStatSet]{player_key: base_period_player_stat_dict}.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).
sample_3pa_overridefloat | NoneNoneWhen given, used as sample3PA instead of sample_team["total_off_3p_attempts"]. Per the upstream comment (LuckUtils.ts:196-198, shared verbatim with calc_def_team_luck_adj's sample_def_3pa_override): "when calc'ing luck on lineups, each lineup gets the total sample as its regression so its average is right over the set" -- i.e. this lets every lineup in a sweep share one common 3PA denominator (the team's) for its regression target, rather than each lineup regressing against its own much smaller, noisier 3PA count. Note that calc_off_player_luck_adj itself does not pass this (its delegation call omits it entirely) -- the jest oracle's own "3P override" cross-check (LuckUtils.test.ts:100-115) instead calls calc_off_team_luck_adj directly with the player's own 3PA as this override, purely to demonstrate the parameter's effect in isolation.
manual_overrideslist[ManualOverride] | NoneNonePer-player 3P%-expectation overrides from the UI. A non-None empty list still activates the team-level override-delta branch (JS array truthiness) -- see the module docstring's landmine note. None (the default) is the "no overrides at all" case.

Returns

An OffLuckAdjustmentDiags dict -- TS-verbatim keys (avgEff, samplePoss, sample3P, sample3PA, base3PA, player3PInfo (per-player detail, sorted by descending shot_info_total_3p), sampleBase3P, regress3P, sampleOff3PRate, sampleOffFGA, sampleOffOrb, sampleOffEfg, sampleOffPpp, sampleDefSos, delta3P, deltaOffEfg, deltaMissesPct, deltaOffPppNoOrb, deltaOffOrbFactor, deltaPtsOffMisses, deltaOffPpp, deltaOffAdjEff).

Example

from sportsdataverse.mbb.mbb_luck import calc_off_team_luck_adj

diags = calc_off_team_luck_adj(
sample_team_on, sample_players_on, base_team, base_players_map, 100.0,
)
print(diags["deltaOffAdjEff"])

# With per-player manual 3P% overrides

diags = calc_off_team_luck_adj(
sample_team_on, sample_players_on, base_team, base_players_map, 100.0,
manual_overrides=[
{"rowId": "Cowan, Anthony", "statName": "off_3p", "newVal": 0.5, "use": True},
],
)

calc_player_weights​

calc_player_weights(ctx: 'RapmPlayerContext') -> 'list[NDArray[np.float64]]'

Build the off/def player-weight (design) matrices for the RAPM solve.

Faithful port of RapmUtils.calcPlayerWeights (RapmUtils.ts:544-595). One row per (filtered) lineup, one column per remaining player; each filled cell is sqrt(lineup_possessions / total_side_possessions) -- the possession-weighted design-matrix entry the ridge regression (Task 3.4) solves against. This is the first function in the module where a dict-shaped RapmPlayerContext gets materialized into a numpy.ndarray -- see the module docstring's "dict -> numpy.ndarray boundary" note.

Parameters

ParameterTypeDefaultDescription
ctxRapmPlayerContextA RapmPlayerContext, e.g. from build_player_context.

Returns

[off_weights, def_weights] -- two numpy.ndarray matrices of shape (num_{off,def}_lineups [+1 if ctx["unbias_weight"] > 0], ctx["num_players"]). The optional extra row (only emitted when ctx["unbias_weight"] > 0 -- always 0.0 in production per build_player_context's hardcoded local, but settable directly on the returned context dict, as the oracle test does) holds each column's unbias_weight-scaled sum-of-squares, an "unbiasing observation" row (RapmUtils.ts:578-593).

Example

from sportsdataverse.mbb.mbb_rapm import calc_player_weights

off_weights, def_weights = calc_player_weights(ctx)
print(off_weights.shape) # (num_off_lineups, num_players)

calc_slow_pseudo_inverse​

calc_slow_pseudo_inverse(player_weight_matrix: 'NDArray[np.float64]', ridge_lambda: 'float', ctx: 'RapmPlayerContext') -> 'NDArray[np.float64]'

Per-parameter variance terms for the ridge-regression standard errors.

Faithful port of the private RapmUtils.calcSlowPseudoInverse (RapmUtils.ts:1544-1557): the same (XᵀX + ridge_lambda·I)⁻¹ as slow_regression's bottomInv, but this function returns the square root of its diagonal instead of the full solver matrix -- the paramErrs term consumed by the standard-error formula (see calculate_sd_rapm).

Parameters

ParameterTypeDefaultDescription
player_weight_matrixNDArray[float64]The off/def design matrix, same shape as slow_regression's.
ridge_lambdafloatThe Tikhonov regularization strength (must match the ridge_lambda used to build the corresponding slow_regression solver, for the SEs to be meaningful).
ctxRapmPlayerContextA RapmPlayerContext -- only ctx["num_players"] is read.

Returns

A length-num_players array, sqrt(diag((XᵀX + λI)⁻¹)).

Example

from sportsdataverse.mbb.mbb_rapm import calc_slow_pseudo_inverse

param_errs = calc_slow_pseudo_inverse(x, 1.0, ctx)