Imputation/SRMI¶
SRMI ¶
Bases: Serializable
Sequential Regression Multiple Imputation (SRMI) class for handling missing data imputation.
This class manages the complete SRMI process including variable setup, model configuration, parallel execution, and result management across multiple implicates and iterations.
Construction groups related settings into sub-objects - see init for the full parameter list:
df, variables, index, imputation_stats : flat, core settings replication : SRMI.Replication - n_implicates / n_iterations / seed parallel : SRMI.Parallel - enabled / variables_per_job / call_inputs / testing storage : SRMI.Storage - path_model / model_name / force_start / save_every_variable / save_every_iteration bootstrap : SRMI.Bootstrap - enabled / index / where defaults : SRMI.Defaults - weight / model / joint / selection / preselection / modeltype / parameters / ordered_categorical (fallback values applied to each Variable added via AddVariable() when that Variable doesn't specify its own)
Raises:
| Type | Description |
|---|---|
Exception
|
If replication.n_implicates < 1 or replication.n_iterations < 1 If path equals path_model_new in load_to_continue_prior |
Examples:
Basic usage:
>>> srmi = SRMI(
... df=data,
... variables=[var1, var2],
... replication=SRMI.Replication(n_implicates=5, n_iterations=10),
... storage=SRMI.Storage(path_model="/path/to/model"),
... )
>>> srmi.run()
With parallel execution:
>>> srmi = SRMI(
... df=data,
... variables=vars_list,
... replication=SRMI.Replication(n_implicates=5, n_iterations=10),
... parallel=SRMI.Parallel(enabled=True, call_inputs=CallInputs(n_cpu=4, mem_in_mb=5000)),
... )
simple_model
classmethod
¶
simple_model(
df: IntoFrameT,
index: list[str] | str | None = None,
variables_to_impute: list[str] | None = None,
classes: dict[str, Class] | None = None,
auto_binary: bool = True,
ordered_categories: dict[str, list] | None = None,
model: dict[Class, ModelType | tuple] | None = None,
yn_pairs: dict[str, str] | None = None,
exclude: dict[str, list[str]] | None = None,
exclude_global: list[str] | None = None,
categorical_predictors: list[str] | None = None,
group_levels: list[str] | str | None = None,
replication: Replication = None,
parallel: Parallel = None,
storage: Storage = None,
bootstrap: Bootstrap = None,
) -> SRMI
survey_kit's equivalent of mice's mice(data, m=5) one-liner
on-ramp - point it at a dataframe and get back a ready-to-.run()
SRMI with sensible per-variable defaults picked for you, that
you can inspect/tweak on .variables before calling .run()
yourself. This is a thin wrapper: all the actual defaulting
logic (which columns need imputing, binary vs. continuous,
ModelType/Parameters per Class, categorical-predictor handling,
yn_pairs via Variable.two_part()) lives in
utilities/auto_detect.py's build_simple_model_variables() - see
its own docstring and this module's docstring for the full
design rationale. Nothing here does anything Variable/
Parameters/SRMI couldn't already do by hand - for anything this
doesn't cover, hand-build a Variable and pass it via
SRMI(variables=[...]) directly instead.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
IntoFrameT
|
The data to impute. |
required |
index
|
list[str] | str | None
|
Same columns you'd pass to SRMI(index=...) directly (a row identifier is auto-generated if you don't give one, same as SRMI's own default) - also auto-excluded from every variable's predictor list here. By default None. |
None
|
variables_to_impute
|
list[str] | None
|
Exactly which columns to impute - if given, no auto-
discovery happens beyond this list (nothing else is scanned
for missingness). If None (the default), every column
(other than |
None
|
classes
|
dict[str, Class] | None
|
{var_name: Variable.Class} - declares a variable's statistical type explicitly. Any imputed variable NOT listed here gets auto-classified as binary (boolean dtype, or numeric with only 0/1 values present) or continuous (auto_binary) - NEVER categorical; a variable that's really ordered_categorical or unordered_categorical always needs an explicit entry here, since category order (or "this numeric-looking column is actually a category code") can't be inferred from the data alone. By default None. |
None
|
auto_binary
|
bool
|
Whether an unclassified variable gets checked for binary-ness
at all (see |
True
|
ordered_categories
|
dict[str, list] | None
|
{var_name: [ordered levels]} - required for any variable classes declares ordered_categorical; raises a clear error if missing. By default None. |
None
|
model
|
dict[Class, ModelType | tuple] | None
|
Per-Class model override. A value can be a bare Variable.ModelType (uses a built-in sensible default Parameters for it) or a (ModelType, parameters_dict) tuple (uses your parameters exactly). Unlisted classes use the built-in defaults: LightGBM for binary/continuous, RandomForest-backed (Multinomial/OrderedCategorical's own default estimator) for both categorical classes. By default None. |
None
|
yn_pairs
|
dict[str, str] | None
|
{value_var: yn_var} - route this variable through Variable.two_part() (semicontinuous/hurdle imputation) instead of a plain single Variable; yn_var must already exist as a column in df. Processed regardless of whether variables_to_impute would otherwise have included value_var/yn_var. By default None. |
None
|
exclude
|
dict[str, list[str]] | None
|
{impute_var: [vars]} - predictor columns to exclude for just that one variable. By default None. |
None
|
exclude_global
|
list[str] | None
|
Columns never used as a predictor for ANY variable built here. By default None. |
None
|
categorical_predictors
|
list[str] | None
|
Predictor columns that must be treated as categorical wherever they're used - passed as categorical_feature=... for a native-categorical modeltype (LightGBM/XGBoost/ CatBoost), or one-hot-encoded via a C(...) formula term otherwise (RandomForest/Multinomial/OrderedCategorical's default estimator/SklearnModel have no native categorical handling). By default None. |
None
|
group_levels
|
list[str] | str | None
|
Passed as group_levels=[...] to every built variable whose modeltype supports it (Regression/RandomForest/XGBoost/ CatBoost/SklearnModel - NOT the binary/continuous default of LightGBM, nor Multinomial/OrderedCategorical - a variable using one of those logs that group_levels was ignored for it). Also auto-folded into exclude_global, so it's never also used as an ordinary predictor. By default None. |
None
|
replication
|
Replication | Parallel | Storage | Bootstrap
|
Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults). |
None
|
parallel
|
Replication | Parallel | Storage | Bootstrap
|
Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults). |
None
|
storage
|
Replication | Parallel | Storage | Bootstrap
|
Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults). |
None
|
bootstrap
|
Replication | Parallel | Storage | Bootstrap
|
Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults). |
None
|
Returns:
| Type | Description |
|---|---|
SRMI
|
Constructed, not yet .run(). |
run ¶
Execute the SRMI imputation process.
Orchestrates the complete imputation workflow including initialization, preprocessing, and execution in parallel or sequential mode.
Notes
The method performs these steps: 1. Creates folders and initializes implicates to be run 2. Preprocesses data (variable selection, hyperparameter tuning) 3. Runs imputation in parallel or sequential mode 4. Saves results and statistics
For parallel execution, creates job files for each iteration and variable subset. For sequential execution, runs each implicate directly.
convergence ¶
Convergence diagnostics for the imputed variables, across implicates and iterations - matches mice's own convergence() function (R/convergence.R) as closely as possible, including the exact algorithm it delegates to for the potential scale reduction factor (rstan::Rhat(), i.e. the rank-normalized, folded, split-Rhat of Vehtari et al. 2021 - see utilities/convergence_diagnostics.py's module docstring for the verified-against-source details).
Callable any time after at least 3 iterations of at least 2 implicates have run (mid-run is fine, not just after a completed SRMI) - it reads whatever each Impute.chain_mean/ chain_std has already accumulated on self.implicates, the same mean/std of each variable's own newly-imputed values every iteration that Impute._post_impute_statistics computes (and that already respects the Variable's own Where restriction, since it's read from df_impute, which Impute.df_impute() has already filtered by Where before this ever sees it).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
diagnostic
|
str
|
"all" (both ac and psrf), "ac" (lag-1 autocorrelation only), or "psrf"/"gr" (potential scale reduction factor only). By default "all". |
'all'
|
parameter
|
str
|
"mean" or "sd" - diagnose the chain means or the chain standard deviations. By default "mean". |
'mean'
|
Returns:
| Type | Description |
|---|---|
IntoFrameT
|
One row per (iteration, variable) - columns ".it", "vrb",
and "ac"/"psrf" per |
plot_convergence ¶
plot_convergence(
parameter: str = "both", path: str | None = None
) -> "plotly.graph_objects.Figure"
Plot the trace lines of the SRMI algorithm - matches mice's own plot(imp) (plot.mids(), R/mids.R) as closely as possible: for each imputed variable, one line per implicate, the chain mean (and/or chain sd) against iteration number. On convergence, the lines within a panel should intermingle and be free of any trend - the same "worm plot" reading as mice's own.
Purely a plotting convenience on top of the same data convergence() reads (Impute.chain_mean/chain_std, accumulated on self.implicates) - there's no "auto-plot" flag anywhere in SRMI's own run() - call this whenever you want, including on an SRMI you've just SRMI.load()ed in a completely different process/environment from the one that ran the imputation (e.g. one where plotly wasn't installed at run time, but is now).
Requires the optional 'plotly' package - not a survey_kit dependency at all (imputation itself never needs it), so nothing about running SRMI requires having it installed; only calling this specific method does.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parameter
|
str
|
"both" (mice's own default - one row of panels for chain means, one for chain sds), "mean", or "sd". By default "both". |
'both'
|
path
|
str | None
|
If given, also save the figure there as a self-contained HTML file (fig.write_html) - no extra dependency beyond plotly itself needed for that, unlike a static image export (which would need kaleido too). By default None (don't save - just return the Figure; display it yourself, e.g. fig.show() in a script or automatically in a notebook). |
None
|
Returns:
| Type | Description |
|---|---|
Figure
|
Always returned (even when |
plot_imputation_quality ¶
plot_imputation_quality(
variable: str | int | list[str | int] | None = None,
kind: str = "density",
sample_k: int | None = None,
seed: int | None = None,
path: str | None = None,
) -> "plotly.graph_objects.Figure"
Plot observed vs. imputed values - a DIFFERENT question from plot_convergence()'s "did the chain stabilize": here it's "do the imputed values look plausible next to the observed ones." Matches mice's own densityplot()/stripplot()/bwplot() as closely as possible (see utilities/quality_diagnostics.py's module docstring for the verified-against-source convention): one group for the truly observed values (pooled once), plus one group per implicate holding ONLY that implicate's own newly imputed values - never the whole column.
Not built here (a materially bigger lift - needs a fitted propensity/detrending model, not just a reshape): mice's propensity-score xyplot() and its own "worm plot" (a detrended Q-Q plot conditional on a covariate - unrelated to the trace lines plot_convergence() draws, despite the similar-sounding name).
No "auto-plot" flag in run() here either, same as plot_convergence() - call this whenever you want, including on an SRMI you've just SRMI.load()ed.
Requires the optional 'plotly' package, same as plot_convergence() - raises a clear ImportError if it isn't installed, only when this is actually called.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
variable
|
str | int | list[str | int] | None
|
Which variable(s) to plot - a name (impute_var), a 0-indexed position in self.variables, a list mixing either, or None for every variable in self.variables. By default None (all). |
None
|
kind
|
str
|
"density" (kernel density per group - numeric/boolean variables only, silently skips a group with fewer than 2 distinct finite values, same wall mice's own densityplot() hits with no workaround), "strip" (every individual point, one column per group), or "box" (five-number-summary box plot per group). By default "density". |
'density'
|
sample_k
|
int | None
|
Only meaningful for kind="strip" - randomly sample at most this many points per (group, variable) to avoid overplotting a large dataset (mice's own guidance: stripplot is best for small datasets, use bwplot/box for large ones - this is the other way to cope with a large one and still see individual points). Ignored for "density"/"box", which should always use every point. By default None (no sampling). |
None
|
seed
|
int | None
|
Seed for the sample_k random sample. By default None. |
None
|
path
|
str | None
|
If given, also save the figure there as a self-contained HTML file. By default None. |
None
|
Returns:
| Type | Description |
|---|---|
Figure
|
|
plot_propensity ¶
plot_propensity(
variable: str | int | list[str | int] | None = None,
kind: str = "density",
n_bins: int = 10,
cv_folds: int = 5,
lightgbm_parameters: dict | None = None,
residual_model_parameters: dict | None = None,
path: str | None = None,
) -> "plotly.graph_objects.Figure"
Response-propensity diagnostic - a different, conditional question from plot_imputation_quality()'s marginal density/ strip/box comparison: instead of "does the imputed marginal distribution look like the observed marginal distribution" (which a correct MAR imputation can legitimately fail, since missing rows can differ systematically on the predictors), this asks "conditional on how similar a row's covariate profile is to a typically-missing row (its response propensity, from a binary LightGBM model of the imputation_flag on that variable's own predictors), does the imputed value track the observed one." See utilities/propensity_diagnostics.py's module docstring for the full rationale, including how kind="density" matches the diagnostic used in Raghunathan & Bondarenko (2007)/Bondarenko & Raghunathan (2016) - and in the user's own SRMI/CPS-ASEC paper (Hokayem, Raghunathan & Rothbaum) - rather than an invented convenience.
No "auto-plot" flag, same as plot_convergence()/ plot_imputation_quality() - call this whenever you want, including on an SRMI you've just SRMI.load()ed. Requires the optional 'plotly' package, same as the other two.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
variable
|
str | int | list[str | int] | None
|
Which variable(s) to plot - a name (impute_var), a 0-indexed position in self.variables, a list mixing either, or None for every variable in self.variables. By default None (all). |
None
|
kind
|
str
|
"density" (default) - regress y on the propensity (a LightGBM regression, one continuous feature, fit on the observed rows only), then compare the KERNEL DENSITY of the residuals - observed vs. each implicate's own imputed rows, scored against that same fit. Similar location AND spread across groups is the good outcome; a shifted or differently-spread residual density for an implicate suggests the imputation model is missing something the missingness mechanism itself depends on. The observed group's own residuals come from cv_folds-way cross- validated (out-of-fold) predictions specifically to avoid biasing them tighter than the (always out-of-sample) implicate residuals - see propensity_diagnostics.propensity_residual_data()'s docstring for why that matters for a single-feature model. "binned_mean" - a simpler, cruder alternative: bin rows by propensity (quantiles of the pooled sample) and compare mean(y), not residuals, within each bin. |
'density'
|
n_bins
|
int
|
Only used by kind="binned_mean" - number of quantile bins of the pooled propensity to group rows into, by default 10. Must be >= 2. |
10
|
cv_folds
|
int
|
Only used by kind="density" - number of cross-validation folds for the observed group's out-of-fold residuals, by default 5. Must be >= 2 to matter; a variable with too few observed rows for the requested fold count (fewer than 2*cv_folds) is silently skipped (logged), same as any other insufficient-data case elsewhere in these diagnostics. |
5
|
lightgbm_parameters
|
dict | None
|
Overrides for the propensity model's LightGBM parameters - merged onto propensity_diagnostics.DEFAULT_PROPENSITY_PARAMETERS (a deliberately shallow/regularized default - an unregularized GBM overfits the in-sample propensity toward 0/1 and collapses the diagnostic). By default None. Each variable's own categorical_feature (from its CatBoost()/ XGBoost()/RandomForest()/SklearnModel()/LightGBM() parameters, whichever it uses) is detected and passed to the propensity model automatically - no need to repeat it here unless you want to override it. |
None
|
residual_model_parameters
|
dict | None
|
Only used by kind="density" - overrides for the y~propensity LightGBM regression's own parameters, merged onto propensity_diagnostics.DEFAULT_RESIDUAL_MODEL_PARAMETERS. By default None. |
None
|
path
|
str | None
|
If given, also save the figure there as a self-contained HTML file. By default None. |
None
|
Returns:
| Type | Description |
|---|---|
Figure
|
|
Variable ¶
Bases: Serializable
Class ¶
Bases: Enum
A variable's statistical TYPE (separate from ModelType, which picks - a specific fitting algorithm for that type). Used by SRMI.simple_model()/utilities/auto_detect.py to pick a sensible default ModelType per variable: binary/continuous can be told apart automatically (dtype/0-1-only check); ordered_categorical and unordered_categorical can never be inferred from the data alone (there's no way to know intended category order, or that a numeric-looking column is really a category code, without being told) - a variable of either categorical kind always needs an explicit Class declaration.
PrePost ¶
Namespace within Variable class for handling pre and post imputation operations.
Currently that can be: 1) a Narwhals Expr (NarwhalsExpression) which is anything you can put in nw.from_native(df).with_columns() 2) a python function handle and parameters which allows you to call an arbitrary function
Predictors ¶
Bases: Serializable
Which predictor variables enter (or are forced into/out of) this variable's model.
Sample ¶
Transforms ¶
Bases: Serializable
Operations to run before/after this variable's imputation each iteration (pre/post), plus once-only bookends around the whole implicate (pre_initialize before the first iteration, post_finalize after the last).
from_legacy
classmethod
¶
from_legacy(
impute_var: str = "",
Where: Expr | None = None,
Where_impute: Expr | None = None,
Where_predict: Expr | None = None,
Where_predict_only_when_not_imputed: bool = False,
bimpute_if_missing: bool = True,
preFunctions=None,
postFunctions=None,
preFunctions_initialize_implicate=None,
predictors_exclude: list = None,
predictors_exclude_first_iteration: list = None,
predictors_require: list = None,
weight: str = "",
joint: dict = None,
header: str = "",
model: str = "",
selection: Selection = None,
preselection: Selection = None,
modeltype: ModelType = None,
modelfunction=None,
parameters: dict = None,
By: list = None,
) -> Variable
Construct a Variable from a flat set of keyword arguments, rather than the grouped sample=Variable.Sample(...)/ transforms=Variable.Transforms(...)/ predictors=Variable.Predictors(...) form Variable() itself takes.
two_part
classmethod
¶
two_part(
df: IntoFrameT,
impute_var: str,
model: list[str] | str,
modeltype: ModelType = ModelType.Regression,
parameters: dict | None = None,
yn_model: list[str] | str | None = None,
yn_modeltype: ModelType | None = None,
yn_parameters: dict | None = None,
yn_var: str | None = None,
yn_missing: Expr | None = None,
value_if_no: float | None = 0,
weight: str = "",
By: list[str] | str | None = None,
) -> tuple[IntoFrameT, list[Variable]]
Shortcut for semicontinuous (point mass at zero + continuous) two-part imputation: a binary y/n ("is impute_var nonzero") variable, imputed first, then impute_var itself, restricted to the y/n==True population - the standard two-part/hurdle approach (as opposed to just PMM/leaf-matching the raw variable directly, which reproduces the point mass for free but assumes a single model/predictor set explains both the participation and intensity margins - see the two_part design discussion this wraps up).
Takes df (needed to derive the y/n column - see below) and
returns (df, variables): df with the y/n column added, and a
plain list, always safe to variables.extend(...) or loop over
- normally [yn_variable, value_variable] (order matters: yn
must precede value in your SRMI variables list, so value sees
this iteration's fresh yn draw, not last iteration's), but
sometimes shorter - see yn_var and the HotDeck/StatMatch note
below. Use the RETURNED df (not your original) to construct
SRMI - SRMI.init needs every impute_var to already exist as
a real column up front, even the y/n one, so it can't be created
lazily via a hook the way the per-iteration consistency fixups
below are.
Mechanics (mostly via existing Variable.Transforms hooks, no per-call special-casing in impute.py): - If creating y/n (yn_var not given): the column is derived once, right here, straight into the returned df (not via pre_initialize - the derivation is a fixed function of df's own observed values, identical for every implicate, so there's nothing to recompute per-implicate): null wherever impute_var is null (or, if yn_missing is given, ALSO wherever that expression is true - additive, never narrower, since a null impute_var can never tell us whether y/n was really True or False), else impute_var != 0. - If yn_var IS given but still has missing values of its own: value_variable's Where (below) would otherwise silently exclude those rows from ever being touched at all (a null never satisfies a boolean filter) - so a real yn_variable still gets built for it, using the same modeltype/ parameters derivation as the create-our-own case, just reading/imputing the caller's own column in place rather than deriving it, and never dropping it at the end (not ours to clean up). If yn_var has no missing values at all, none of this applies - value_variable alone is returned. - Whenever a real yn_variable is built (either case above) and its own donation is genuine (pmm/leaf, or HotDeck/ StatMatch - always donation-based), impute_var itself rides along in its donate_list - so a row whose y/n was just resolved gets its value from that SAME matched donor, rather than from a possibly different one value_variable's own later pass would find. A harmless no-op when yn's error draw is Random instead, which never donates anything. - value_variable.sample.Where restricts value's donor pool AND recipients to yn==True rows. - value_variable.transforms.pre runs every iteration, before value's own fit/donation: wherever yn is currently True but value == 0 (stale from a prior iteration when yn was False), null it out (so it's a genuine recipient again this iteration); wherever yn is currently False, force value to value_if_no. - yn_variable.transforms.post_finalize drops the scratch y/n column once, after the implicate's last iteration - only when this function created it (see yn_var).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
IntoFrameT
|
Source data - read to derive the y/n column, and to return the augmented copy you should actually build SRMI from. |
required |
impute_var
|
str
|
The semicontinuous target. |
required |
model
|
list[str] | str
|
Predictors for impute_var (value). Also yn's predictors, unless yn_model overrides them. |
required |
modeltype
|
ModelType
|
value's modeltype. Supported: Regression, pmm, LightGBM (fit a real Logit/binary-objective variant for yn); RandomForest, XGBoost, CatBoost, SklearnModel (yn reuses the SAME estimator factory via ModelType.OrderedCategorical with categories=[False, True] - no regressor/classifier estimator swap, which isn't generically possible: sklearn/ XGBoost/CatBoost each have model-family-specific hyperparameters, like criterion/objective/loss_function, that don't transfer across that boundary, and there's no way to do it at all for SklearnModel's arbitrary factory); HotDeck, StatMatch (donation doesn't care about target dtype, no swap needed - see the collapse behavior below). Anything else (Multinomial, OrderedCategorical) raises - not semicontinuous-shaped as value's own type. By default Variable.ModelType.Regression. |
Regression
|
parameters
|
dict
|
value's parameters (e.g. Parameters.RandomForest(...)). By default None ({}). |
None
|
yn_model
|
list[str] | str | None
|
Predictors for yn, if they should differ from model - the participation and intensity margins often don't share predictors/mechanism. By default None (same as model). |
None
|
yn_modeltype
|
ModelType | None
|
Explicit override for yn's modeltype - pass together with yn_parameters for full manual control. By default None (derived from modeltype per the modeltype docstring above). |
None
|
yn_parameters
|
dict | None
|
Explicit override for yn's parameters. If given without yn_modeltype, yn_modeltype defaults to modeltype (or OrderedCategorical, if modeltype is one of the four that reuse it) rather than being derived further - if you're supplying parameters yourself, supply the modeltype too if it's not that default. By default None (derived). |
None
|
yn_var
|
str | None
|
Reuse an existing y/n column instead of deriving one from impute_var. If it's already fully observed, this function builds ONLY value_variable (a 1-item list), restricted to yn_var==True. If it still has missing values of its own, this function ALSO builds a yn_variable to resolve them (using the same modeltype/yn_modeltype/yn_parameters derivation as the create-our-own case - see the mechanics note above), so the returned list is still [yn_variable, value_variable] in that case - it's only "your own concern" when there's genuinely nothing left for it to do. Never dropped at the end either way - it's your column. By default None (create one, named f"{imputevar}_yn_"). |
None
|
yn_missing
|
Expr | None
|
Only meaningful when yn_var is not given. Rows where this is true are ALSO treated as yn-missing, on top of impute_var.is_null() (additive, not a replacement - see the mechanics note above for why). By default None. |
None
|
value_if_no
|
float | None
|
What value becomes on yn==False rows, every iteration. By default 0 (pass None to leave it null instead). |
0
|
weight
|
str
|
Applied identically to both variables. By default "". |
''
|
By
|
list[str] | str | None
|
Applied identically to both variables. By default None. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[IntoFrameT, list[Variable]]
|
(df, variables) - df augmented with the y/n column (only when this function created one - unchanged otherwise); variables is [yn_variable, value_variable] normally, [value_variable] alone only when either (a) yn_var was given and is already fully observed (nothing left for a yn_variable to do), or (b) modeltype is HotDeck/StatMatch and there's no signal at all (no yn_var, yn_missing, yn_model, yn_modeltype, or yn_parameters) that the two margins should be modeled differently - a single donation pass on value, unrestricted, already reproduces the point mass for free in that case, so a redundant yn model is skipped (with a warning). |
union_required_predictors ¶
Add any predictors_require variables not already present in formula.
Used after variable selection (LASSO/stepwise) replaces the model formula, to guarantee required predictors survive selection - the selection methods themselves never see predictors_require (they don't receive the Variable object), so this has to happen here, in the caller, once selection has returned.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
formula
|
str
|
A model formula, typically the one returned by Selection.run()/lasso()/stepwise(). |
required |
Returns:
| Type | Description |
|---|---|
str
|
formula, with any missing required predictors added. |
Parameters ¶
Factory class for creating parameter dictionaries for different imputation methods.
Provides static methods to generate properly formatted parameter dictionaries with validation and default values for each imputation approach.
ErrorDraw ¶
Bases: Enum
Random = 0 pmm = 1 leaf = 2
CatBoost
staticmethod
¶
CatBoost(
parameters: dict | None = None,
error: ErrorDraw = ErrorDraw.pmm,
random_share: float = 1.0,
cv_folds: int = 0,
categorical_feature: list[str] | str | None = None,
group_levels: list[str] | None = None,
group_shrinkage_k: float = 10.0,
parameters_pmm: dict = None,
tune: bool = False,
tuner=None,
) -> dict
Parameters for CatBoost-based imputation (mean regression only).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parameters
|
dict
|
Keyword arguments passed straight through to catboost.CatBoostRegressor(**parameters), by default None (CatBoostRegressor's own defaults, plus verbose=0). |
None
|
error
|
ErrorDraw
|
Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring. |
pmm
|
random_share
|
float
|
Fraction of data to use for fitting, by default 1.0. |
1.0
|
cv_folds
|
int
|
See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off). |
0
|
categorical_feature
|
list[str] | str | None
|
Predictor column names to treat as native categoricals (CatBoost's own ordered-target-statistic categorical handling, not one-hot encoding) - see XGBoost()'s categorical_feature docstring for why this is separate from a formula's C(...) syntax. By default None (no native categoricals). |
None
|
group_levels
|
list[str] | None
|
See Regression()'s group_levels docstring, by default None (off). |
None
|
group_shrinkage_k
|
float
|
See Regression()'s group_shrinkage_k docstring, by default 10.0. |
10.0
|
parameters_pmm
|
dict
|
PMM parameters if using PMM error drawing, by default None. |
None
|
tune
|
bool
|
Whether to tune hyperparameters for this run, by default False.
No effect if |
False
|
tuner
|
Tuner
|
Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None. |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
CatBoost parameter dictionary |
HotDeck
staticmethod
¶
HotDeck(
model_list: list[str] | list[list[str]] = None,
donate_list: list | None = None,
n_hotdeck_array: int = 3,
sequential_drop: bool = True,
) -> dict
Parameters for hot HotDeck imputation
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_list
|
list[str] | list[list[str]]
|
Each model is a list of variables that are used as match keys. model_list can either be a list of strings (the model itself) or it can be a list of lists of strings (sequential hot deck to match on) |
None
|
donate_list
|
list
|
Additional variables to impute together, by default None I.e., you can predict earnings amount to find a donor, but then also impute hours worked and weeks worked with it. |
None
|
n_hotdeck_array
|
int
|
Size of hot deck donor arrays, by default 3 |
3
|
sequential_drop
|
bool
|
Drop variables sequentially until matches found, by default True If model_list is a list of strings (one model), should we sequentially drop the last variable until all recipients find a donor? Makes it easier to set the hot deck/stat match up. Whatever recipients are STILL unmatched after every model in model_list (including every dropped-down level of this cascade) has been tried get matched fully at random instead, with a logged warning, rather than being left unmatched forever - see impute.py's hotdeck()/statmatch(), the guaranteed-to-succeed last resort applies regardless of sequential_drop or what model_list contains. |
True
|
Returns:
| Type | Description |
|---|---|
dict
|
Hot deck parameter dictionary |
LightGBM
staticmethod
¶
LightGBM(
tune: bool = False,
parameters: dict | None = None,
tuner=None,
quantiles: list = None,
error: ErrorDraw = ErrorDraw.pmm,
cv_folds: int = 0,
parameters_pmm: dict | None = None,
) -> dict
Parameters for LightGBM-based imputation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tune
|
bool
|
Whether to tune hyperparameters for this run, by default False.
No effect if |
False
|
parameters
|
dict | None
|
LightGBM model parameters, by default None |
None
|
tuner
|
Tuner
|
Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes - see utilities.tuning.Tuner / HyperparameterSpace), by default None |
None
|
quantiles
|
list
|
Quantiles for quantile regression, by default None |
None
|
error
|
ErrorDraw
|
How to convert yhat from LGBM into imputes. If pmm, draw from nearest yhat neighbors, for example. The default is ErrorDraw.pmm. |
pmm
|
cv_folds
|
int
|
See Parameters._tabular_ml_params's cv_folds docstring - same tabular-ML cv_folds mechanism RandomForest()/XGBoost()/ CatBoost()/SklearnModel() use, by default 0 (off). |
0
|
parameters_pmm
|
dict | None
|
PMM parameters if using PMM error drawing, by default None. "winsor" is dropped even if present - unlike pmm()/Regression(), LightGBM's own fit path never winsorizes impute_var before training (that's only done by _pmm_winsorize_for_fit, called from pmm()/regression() - the latter also covers RandomForest/ XGBoost/CatBoost/SklearnModel, which route through regression(), but LightGBM has its own separate fit path that doesn't). |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
LightGBM parameter dictionary |
Raises:
| Type | Description |
|---|---|
Exception
|
If quantiles are not between 0 and 1 |
Multinomial
staticmethod
¶
Multinomial(
parameters: dict | None = None,
donate_list: list[str] | None = None,
donate_by: list[str] | str | None = None,
random_share: float = 1.0,
) -> dict
Parameters for imputing an unordered categorical variable with 3+ levels via a RandomForestClassifier and donor matching on leaf co-occurrence - see imputation/utilities/leaf_donor_matching.leaf_cooccurrence_match for the donor-selection mechanism (the same one mice's rf method uses: pool donors sharing a leaf with the recipient across every tree, draw one uniformly at random).
This is a distinct imputation shape from RandomForest()/XGBoost()/CatBoost()/SklearnModel() above - those are all mean regression (a single continuous yhat, optionally PMM-matched); this is genuinely multi-class classification, with the donor match itself driven by which trees group two rows together, not by distance on a predicted scalar. There's no error=/cv_folds= here for that reason - PMM-style matching on a scalar prediction and "refit on folds to correct in-sample bias" both assume a scalar yhat, which doesn't exist for this method.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parameters
|
dict
|
Keyword arguments passed straight through to sklearn.ensemble.RandomForestClassifier(**parameters), by default None (RandomForestClassifier's own defaults). No native categorical predictor handling (same restriction as RandomForest()) - model= can be a plain column list or an R-style formula string; either way, a categorical predictor needs to end up numeric before reaching this model, via the formula's own C(...) one-hot encoding or, for the list form, by already being numeric-coded. |
None
|
donate_list
|
list[str]
|
Additional variables to impute together from the same matched donor, by default None. |
None
|
donate_by
|
list[str] | str | None
|
Grouping variable(s) - donors are only matched within the recipient's own group, by default None. |
None
|
random_share
|
float
|
Fraction of data to use for fitting, by default 1.0. |
1.0
|
Returns:
| Type | Description |
|---|---|
dict
|
Multinomial parameter dictionary |
NearestNeighbor
staticmethod
¶
NearestNeighbor(
match_to: str | list[str],
logit: bool = False,
parameters_pmm: dict = None,
) -> dict
Convenience builder for nearest-neighbor PMM matching, purely for readability/discoverability - there's no dedicated Variable.ModelType.NearestNeighbor; this just returns Parameters.Regression()'s own dict shape, so use it with Variable.ModelType.Regression. Matching on a fitted OLS/Logit prediction (a principled notion of "similar") is strictly better than matching on raw, unweighted, unlearned multivariate distance across match_to's predictors directly, so this always routes through Regression rather than a distance-based implementation of its own.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
match_to
|
str | list[str]
|
The predictor(s) to match on. NOT stored anywhere in the
returned dict - Parameters.XXX() functions never set
Variable.model themselves (true of every one of them, not
specific to this one), so pass this SAME list as the
Variable's own |
required |
logit
|
bool
|
impute_var is binary - fit Logit instead of OLS (still matched via PMM on the fitted probability, same as OLS's fitted mean). By default False (OLS). |
False
|
parameters_pmm
|
dict
|
knearest/donate_list/donate_by/winsor - same as pmm()/ Regression(). By default None (Parameters.pmm()'s own defaults). |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
Parameters.Regression(model=OLS or Logit, error=pmm, ...)'s parameter dictionary. |
OrderedCategorical
staticmethod
¶
OrderedCategorical(
categories: list,
parameters: dict | None = None,
estimator: Callable[[], object] | None = None,
error: ErrorDraw = ErrorDraw.pmm,
random_share: float = 1.0,
categorical_feature: list[str] | str | None = None,
estimator_prepare_data: Callable[
[object, object], tuple[object, object]
]
| None = None,
donate_list: list[str] | None = None,
donate_by: list[str] | str | None = None,
knearest: int = 10,
) -> dict
Parameters for imputing an ORDERED categorical variable (e.g. an
education level or a Likert scale) - unlike Multinomial()
(unordered, classification), this fits a mean-regression
estimator against an integer rank encoding of categories
(lowest/coarsest to highest/finest), then donates the REAL
observed category from a matched donor - never the numeric rank,
and never a category that wasn't actually observed, the same
guarantee Multinomial() and PMM/leaf donor matching already give
elsewhere.
This reuses the same donor-matching machinery RandomForest()/ XGBoost()/CatBoost() etc. use for their own error=pmm/leaf, just matched on the predicted rank instead of impute_var's own (non-numeric) value - see impute.py's ordered_categorical().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
categories
|
list
|
Every observed value of impute_var, ordered from lowest/coarsest to highest/finest (e.g. ["less_than_hs", "hs_grad", "some_college", "college_grad"]). Imputation raises if any observed value isn't in this list. |
required |
parameters
|
dict
|
Keyword arguments for the default estimator
(sklearn.ensemble.RandomForestRegressor(**parameters)) -
ignored if |
None
|
estimator
|
Callable[[], object]
|
Zero-arg factory for a fresh, unfitted mean-regression
estimator (e.g. |
None
|
error
|
ErrorDraw
|
pmm (knearest on the predicted rank) or leaf (tree leaf
co-occurrence - needs |
pmm
|
random_share
|
float
|
Fraction of data to use for fitting, by default 1.0. |
1.0
|
categorical_feature
|
list[str] | str | None
|
Predictor columns to cast to a fixed-category dtype before fitting, for an estimator with native categorical handling (XGBoost/CatBoost) - see RandomForest()/XGBoost()'s docstring. By default None. |
None
|
estimator_prepare_data
|
Callable[[df_model, df_impute], (df_model, df_impute)]
|
See SklearnModel()'s prepare_data docstring. By default None. |
None
|
donate_list
|
list[str]
|
Additional variables to impute together from the same matched donor, by default None. |
None
|
donate_by
|
list[str] | str | None
|
Grouping variable(s) - donors are only matched within the recipient's own group, by default None. |
None
|
knearest
|
int
|
Only used when error=pmm - number of nearest neighbors (on predicted rank) to draw from, by default 10. |
10
|
Returns:
| Type | Description |
|---|---|
dict
|
OrderedCategorical parameter dictionary |
RandomForest
staticmethod
¶
RandomForest(
parameters: dict | None = None,
error: ErrorDraw = ErrorDraw.pmm,
random_share: float = 1.0,
cv_folds: int = 0,
group_levels: list[str] | None = None,
group_shrinkage_k: float = 10.0,
parameters_pmm: dict = None,
tune: bool = False,
tuner=None,
) -> dict
Parameters for RandomForestRegressor-based imputation (mean regression only - scikit-learn's RandomForestRegressor has no native quantile-loss support). No native categorical handling either - categorical predictors need to be encoded (e.g. via a "~...+C(var)+..." formula) before reaching this model, same as OLS/Logit.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parameters
|
dict
|
Keyword arguments passed straight through to sklearn.ensemble.RandomForestRegressor(**parameters), by default None (RandomForestRegressor's own defaults). |
None
|
error
|
ErrorDraw
|
Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring. |
pmm
|
random_share
|
float
|
Fraction of data to use for fitting, by default 1.0. |
1.0
|
cv_folds
|
int
|
See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off). |
0
|
group_levels
|
list[str] | None
|
See Regression()'s group_levels docstring, by default None (off). |
None
|
group_shrinkage_k
|
float
|
See Regression()'s group_shrinkage_k docstring, by default 10.0. |
10.0
|
parameters_pmm
|
dict
|
PMM parameters if using PMM error drawing, by default None. |
None
|
tune
|
bool
|
Whether to tune hyperparameters for this run, by default False.
No effect if |
False
|
tuner
|
Tuner
|
Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None. |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
RandomForest parameter dictionary |
Regression
staticmethod
¶
Regression(
model: RegressionModel = RegressionModel.OLS,
error: ErrorDraw = ErrorDraw.pmm,
random_share: float = 1.0,
group_levels: list[str] | None = None,
group_shrinkage_k: float = 10.0,
parameters_pmm: dict = None,
) -> dict
Parameters for regression-based imputation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
RegressionModel
|
Type of regression model, by default RegressionModel.OLS |
OLS
|
error
|
ErrorDraw
|
Method for drawing errors, by default ErrorDraw.Random. ErrorDraw.leaf isn't usable here (OLS/Logit have no tree structure to match donors on) - see RandomForest()/ XGBoost()/CatBoost()/SklearnModel() for that. If pmm, draw from nearest yhat neighbors, for example. If Random, draw from observed errors for modeled observations |
pmm
|
random_share
|
float
|
Fraction of data to use for regression, by default 1.0 Use less memory by running the regression on a random subset? |
1.0
|
group_levels
|
list[str] | None
|
Nested grouping columns, COARSEST to FINEST (e.g. ["state", "county", "hhid"]), for a cheap shrinkage-heuristic stand-in for a random-intercept term - not a real mixed model, see Impute._nested_group_shrinkage's docstring for the exact mechanism and its limits. Re-estimated fresh from each SRMI iteration's fit residuals and persisted into the working data (as a plain column, the same way donate_list values ride along) for the next iteration to build on - no separate inner convergence loop. By default None (off). |
None
|
group_shrinkage_k
|
float
|
Shrinkage constant for group_levels - a group needs roughly this many observations before its own mean starts to dominate over being pulled toward 0. Only meaningful when group_levels is set. By default 10.0. |
10.0
|
parameters_pmm
|
dict
|
PMM parameters if using PMM error drawing, by default None |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
Regression parameter dictionary |
SklearnModel
staticmethod
¶
SklearnModel(
factory: Callable[[], object],
error: ErrorDraw = ErrorDraw.pmm,
random_share: float = 1.0,
cv_folds: int = 0,
categorical_feature: list[str] | str | None = None,
prepare_data: Callable[
[object, object], tuple[object, object]
]
| None = None,
group_levels: list[str] | None = None,
group_shrinkage_k: float = 10.0,
parameters_pmm: dict = None,
tune: bool = False,
tuner=None,
) -> dict
Parameters for imputation using any sklearn-compatible estimator you bring yourself (mean regression only) - the escape hatch for anything without a dedicated RandomForest()/XGBoost()/CatBoost() function.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
factory
|
Callable[[], object]
|
Zero-arg callable returning a fresh, unfitted
sklearn-compatible estimator (supporting .fit(X, y) and
.predict(X), plus .fit(..., sample_weight=...) if you're using
a weight variable). E.g. |
required |
error
|
ErrorDraw
|
Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring. |
pmm
|
random_share
|
float
|
Fraction of data to use for fitting, by default 1.0. |
1.0
|
cv_folds
|
int
|
See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off). |
0
|
categorical_feature
|
list[str] | str | None
|
Predictor column names to cast to a fixed-category dtype
(polars Enum) in the model matrix before fitting - useful if
your own estimator wants that dtype for native categorical
handling, the same way XGBoost/CatBoost do. This only casts
the dtype; if your model needs the categorical columns
communicated some other way too (e.g. a constructor kwarg
naming them), use |
None
|
prepare_data
|
Callable[[df_model, df_impute], (df_model, df_impute)]
|
Full control over preparing the model matrix right before fitting/predicting - called once, on both the donor pool and recipient frames together (so anything derived from both, like fixed Enum categories, stays consistent between them), after the regular formula/model-matrix construction and before your factory's model is fit. By default None (no extra prep beyond categorical_feature, if any). |
None
|
group_levels
|
list[str] | None
|
See Regression()'s group_levels docstring, by default None (off). |
None
|
group_shrinkage_k
|
float
|
See Regression()'s group_shrinkage_k docstring, by default 10.0. |
10.0
|
parameters_pmm
|
dict
|
PMM parameters if using PMM error drawing, by default None. |
None
|
tune
|
bool
|
Whether to tune hyperparameters for this run, by default False.
No effect if |
False
|
tuner
|
Tuner
|
Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None. |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
Custom sklearn-model parameter dictionary |
StatMatch
staticmethod
¶
Parameters for hot statistical match imputation
Stat match and hot deck are basically the same, but the hot deck iterates over the data carrying arrays of possible donor values whereas the stat match just does a join of donors and recipients
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_list
|
list[str] | list[list[str]]
|
Each model is a list of variables that are used as match keys. model_list can either be a list of strings (the model itself) or it can be a list of lists of strings (sequential hot deck to match on) |
None
|
donate_list
|
list
|
Additional variables to impute together, by default None I.e., you can predict earnings amount to find a donor, but then also impute hours worked and weeks worked with it. |
None
|
sequential_drop
|
bool
|
Drop variables sequentially until matches found, by default False If model_list is a list of strings (one model), should we sequentially drop the last variable until all recipients find a donor? Makes it easier to set the hot deck/stat match up. |
False
|
Returns:
| Type | Description |
|---|---|
dict
|
stat match parameter dictionary |
TwoSampleRegression
staticmethod
¶
TwoSampleRegression(
model: "Parameters.RegressionModel" = None,
is_boolean: bool = False,
bins: int = 10,
bin_by: list[str] | str | None = None,
percentile_cuts: list[float] | None = None,
save_percentile_cuts: bool = False,
round_impute_var_digits: int = 4,
continuous_qtiles_y_cuts: list[float] | None = None,
continuous_qtiles_interpolate_by_bin: bool = False,
min_n_x_var: int = 0,
draw_error: bool = False,
random_share: float = 1.0,
save_disclosure_support: bool = False,
path_save: str = "",
path_load: str = "",
load_from_save: bool = False,
) -> dict
Parameters for two-sample regression imputation.
Fits a regression (OLS/Logit) on one sample, then imputes a (possibly different) sample by binning the predicted yhat into percentile groups and drawing from the EMPIRICAL distribution of y within each bin - never a value predicted directly by the model. Unlike pmm/leaf error draws, nothing is donated from one recipient row to another; every draw comes from the model sample's own observed y values (or residuals, if draw_error=True), aggregated per bin.
The fitted model (coefficients, bin cutoffs, per-bin distribution) can be persisted to disk (path_save) and reloaded later (path_load/load_from_save) to impute a genuinely separate sample that was never in memory at the same time as the model sample - the "two-sample" in the name. Persistence is plain CSV/JSON files in a variable-specific folder (never pickle - not reliably portable across machines/environments/library versions - and never a bundled archive either, so every file can be opened directly without an extraction step), so model= must resolve to numeric-only predictors (a plain column list, or a formula with no C(...)/factor terms) - a fitted categorical encoding has no portable plain-text representation. Pre-encode any categorical predictor into numeric/dummy columns yourself first.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
RegressionModel
|
OLS or Logit, by default RegressionModel.OLS (Logit if is_boolean=True and model isn't passed explicitly). |
None
|
is_boolean
|
bool
|
Whether impute_var is binary, by default False. Controls both the default model choice and the draw mechanism: bin-level P(y=1) + a Bernoulli draw, instead of a bin-level empirical quantile distribution. |
False
|
bins
|
int
|
Number of percentile bins to split predicted yhat into, by default 10 - each bin gets its own P(y=1) (boolean) or empirical quantile distribution (continuous) to draw from. |
10
|
bin_by
|
list[str] | str | None
|
Additional grouping columns - compute a separate distribution per (bin_by, prediction bin) cell instead of per prediction bin alone, by default None. |
None
|
percentile_cuts
|
list[float] | None
|
Explicit bin cut points (0-100 scale) instead of |
None
|
save_percentile_cuts
|
bool
|
Freeze the bin cutoffs the first time they're computed under path_save, reusing the same cutoffs on every later call (e.g. across SRMI's repeated iterations) instead of recomputing them fresh each time - keeps disclosure-reviewed bin boundaries stable while the regression and distribution still refit every call. Independent of load_from_save, which freezes everything. By default False. |
False
|
round_impute_var_digits
|
int
|
Digits to DRB-round (see utilities.rounding.drb_round_table) the predicted score and impute_var to before computing bin cutoffs/distributions and before persisting them, by default 4. |
4
|
continuous_qtiles_y_cuts
|
list[float] | None
|
Quantile levels (0-1 scale) of the empirical y (or residual) distribution to compute per bin, later interpolated between to draw a value - see utilities.draw_from_quantiles. By default [0.1, 0.25, 0.5, 0.75, 0.9] if None. |
None
|
continuous_qtiles_interpolate_by_bin
|
bool
|
Compute the Census-style interpolation interval (Statistics(quantile_interpolated=True)) separately within each bin instead of once globally, by default False - more accurate when bins vary a lot in spread, at the cost of time. |
False
|
min_n_x_var
|
int
|
Minimum non-zero observations required per predictor to keep it in the model, by default 0 (no restriction) - for disclosure, same as _build_model_matrix's min_n_x_var. |
0
|
draw_error
|
bool
|
Draw from the bin's empirical distribution of residuals (yhat - y) and add the draw to yhat, instead of drawing directly from the bin's empirical distribution of y itself. Only meaningful when is_boolean=False. By default False. |
False
|
random_share
|
float
|
Fraction of the model sample to fit on, by default 1.0. |
1.0
|
save_disclosure_support
|
bool
|
Also persist a disclosure-review audit trail alongside the model (only when path_save is set): the exact record count in every (bin_by, prediction bin) cell the distribution was built from, and how many model+recipient records land exactly on each bin cutoff (relevant when the underlying data has heaping/rounding). Never affects the fitted model, the cutoffs, or the draw itself - purely an extra file for a human reviewer to check cell sizes with. By default False - these are real record counts, so leave this off unless someone is actually going to review them. |
False
|
path_save
|
str
|
Directory to persist the fitted model/cutoffs/distribution to (as plain files under f"{path_save}/{impute_var}/"), by default "" (don't persist). |
''
|
path_load
|
str
|
Directory to load a previously-persisted model from when load_from_save=True, by default "" (falls back to path_save). |
''
|
load_from_save
|
bool
|
Skip fitting entirely and impute purely from a previously persisted model (path_load/path_save) - the model sample doesn't need to be present in this run at all. By default False. |
False
|
Returns:
| Type | Description |
|---|---|
dict
|
Two-sample regression parameter dictionary |
XGBoost
staticmethod
¶
XGBoost(
parameters: dict | None = None,
error: ErrorDraw = ErrorDraw.pmm,
random_share: float = 1.0,
cv_folds: int = 0,
categorical_feature: list[str] | str | None = None,
group_levels: list[str] | None = None,
group_shrinkage_k: float = 10.0,
parameters_pmm: dict = None,
tune: bool = False,
tuner=None,
) -> dict
Parameters for XGBoost-based imputation (mean regression only - see Parameters._tabular_ml_params's cv_folds docstring for why an honest donor-pool prediction matters more for flexible models like this one).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parameters
|
dict
|
Keyword arguments passed straight through to xgboost.XGBRegressor(**parameters), by default None (XGBRegressor's own defaults). |
None
|
error
|
ErrorDraw
|
Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring. |
pmm
|
random_share
|
float
|
Fraction of data to use for fitting, by default 1.0. |
1.0
|
cv_folds
|
int
|
See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off). |
0
|
categorical_feature
|
list[str] | str | None
|
Predictor column names to treat as native categoricals
(XGBoost's own histogram-based categorical splits, not
one-hot encoding). Works with either form of the Variable's
|
None
|
group_levels
|
list[str] | None
|
See Regression()'s group_levels docstring, by default None (off). |
None
|
group_shrinkage_k
|
float
|
See Regression()'s group_shrinkage_k docstring, by default 10.0. |
10.0
|
parameters_pmm
|
dict
|
PMM parameters if using PMM error drawing, by default None. |
None
|
tune
|
bool
|
Whether to tune hyperparameters for this run, by default False.
No effect if |
False
|
tuner
|
Tuner
|
Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None. |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
XGBoost parameter dictionary |
pmm
staticmethod
¶
pmm(
knearest: int = 10,
model: RegressionModel = RegressionModel.OLS,
donate_list: list[str] = None,
winsor: tuple[float, float] = [0, 1],
donate_by: list[str] | str | None = None,
) -> dict
Parameters for predictive mean matching imputation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
knearest
|
int
|
Number of nearest neighbors for matching, by default 10 |
10
|
model
|
RegressionModel
|
Regression model type, by default RegressionModel.OLS |
OLS
|
donate_list
|
list[str]
|
Additional variables to impute together, by default None i.e., you can predict earnings amount to find a donor, but then also impute hours worked and weeks worked with it. |
None
|
winsor
|
tuple[float, float]
|
Winsorization percentiles, by default [0, 1] |
[0, 1]
|
donate_by
|
list[str] | str | None
|
Grouping variables for donation, by default None |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
PMM parameter dictionary |
Selection ¶
Bases: Serializable
Handles variable selection for imputation models.
Supports various selection methods including LASSO, stepwise selection, and custom functions to reduce model dimensionality.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
method
|
Method
|
Selection method to use, by default Method.No |
No
|
parameters
|
dict
|
Method-specific parameters, by default None |
None
|
select_within_by
|
bool
|
Whether to run selection within each by-group, by default True |
True
|
function
|
callable
|
Custom selection function, by default None |
None
|
Source code in src/survey_kit/imputation/selection.py
18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 | |
Parameters ¶
Source code in src/survey_kit/imputation/selection.py
46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 | |
lasso ¶
lasso(
nfolds: int = 5,
type_measure: str = "default",
include_base_with_interaction: bool = True,
winsorize: tuple[float, float] | None = None,
continuous: bool = False,
binomial: bool = False,
missing_dummies: bool = True,
optimal_lambda: float = None,
optimal_lambda_from_pre: bool = True,
scale_lambda: float = 1.0,
) -> dict
Parameters for LASSO variable selection method.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
nfolds
|
int
|
Number of folds for cross-validation during LASSO regression. Used to determine optimal lambda value through k-fold cross-validation. |
5
|
type_measure
|
str
|
Type of measure to use for cross-validation error. Determines how model performance is evaluated during lambda selection. Options are "default", "mse", "deviance", "class", "auc", "mae". |
"default"
|
include_base_with_interaction
|
bool
|
Whether to include base variables when interaction terms are selected. If True, base variables are automatically included when their interactions are selected by LASSO. |
True
|
winsorize
|
tuple[float, float] or None
|
Percentile bounds for winsorizing the dependent variable. Tuple of (low_percentile, high_percentile) to cap extreme values. None means no winsorization. |
None
|
continuous
|
bool
|
Force treatment of dependent variable as continuous, overriding automatic detection. |
False
|
binomial
|
bool
|
Force treatment of dependent variable as binomial, overriding automatic detection. |
False
|
missing_dummies
|
bool
|
Whether to create dummy variables for missing values in predictor variables. |
True
|
optimal_lambda
|
float or None
|
Pre-specified optimal lambda value for LASSO regularization. If None, optimal lambda will be determined through cross-validation. |
None
|
optimal_lambda_from_pre
|
bool
|
Whether to use optimal lambda from a previous preselection step. |
True
|
scale_lambda
|
float
|
Scaling factor applied to the optimal lambda value. Values > 1 make regularization stronger (fewer variables selected), values < 1 make it weaker (more variables selected). |
1.0
|
Returns:
| Type | Description |
|---|---|
dict
|
Dictionary containing all parameter values for LASSO selection. |
Raises:
| Type | Description |
|---|---|
Exception
|
When type_measure is not one of the acceptable values. |
Source code in src/survey_kit/imputation/selection.py
stepwise ¶
stepwise(
nfolds: int = 5,
scoring: str = "neg_mean_squared_error",
include_base_with_interaction: bool = True,
winsorize: tuple[float, float] | None = None,
missing_dummies: bool = True,
min_features_to_select: int = 10,
)
Parameters for stepwise variable selection method using Recursive Feature Elimination with Cross-Validation (RFECV).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
nfolds
|
int
|
Number of folds for cross-validation during stepwise selection. Used in RFECV to evaluate feature importance. |
5
|
scoring
|
str
|
Scoring metric used to evaluate model performance during cross-validation. Should be a valid sklearn scoring parameter. |
"neg_mean_squared_error"
|
include_base_with_interaction
|
bool
|
Whether to include base variables when interaction terms are selected. If True, base variables are automatically included when their interactions are selected. |
True
|
winsorize
|
tuple[float, float] or None
|
Percentile bounds for winsorizing the dependent variable. Tuple of (low_percentile, high_percentile) to cap extreme values. None means no winsorization. |
None
|
missing_dummies
|
bool
|
Whether to create dummy variables for missing values in predictor variables. |
True
|
min_features_to_select
|
int
|
Minimum number of features that must be selected by the stepwise procedure. Prevents over-reduction of the feature set. |
10
|
Returns:
| Type | Description |
|---|---|
dict
|
Dictionary containing all parameter values for stepwise selection. |
Source code in src/survey_kit/imputation/selection.py
tuning ¶
Categorical
dataclass
¶
A hyperparameter chosen from a fixed set of values.
FloatRange
dataclass
¶
A float hyperparameter, uniform (or log-uniform) between low/high.
HyperparameterSpace ¶
A named, typed hyperparameter search space - replaces a bare {"num_leaves": [2, 256]} dict with self-documenting IntRange/FloatRange/ Categorical fields, e.g.:
HyperparameterSpace(
num_leaves=IntRange(2, 256),
learning_rate=FloatRange(1e-3, 0.3, log=True),
boosting=Categorical(["gbdt", "dart"]),
)
Extensible: any object with a .suggest(trial, name) method can be used
as a field value - a new distribution type is just a new small class,
not a change to HyperparameterSpace or Tuner.
IntRange
dataclass
¶
An integer hyperparameter, uniform (or log-uniform) between low/high.
Objective ¶
Bases: Enum
A scoring function for comparing a model's raw predictions against
holdout labels - used by Tuner.run_lightgbm (LightGBM's own .predict()
doesn't go through an sklearn Estimator, so it can't use run_estimator's
sklearn scoring= string instead). Values must NOT be plain functions -
Enum silently treats function-valued (or otherwise descriptor-like, e.g.
functools.partial on newer Python) class attributes as methods rather
than registering them as real members, so string values are used here
instead and the actual scoring function is looked up via
_objective_scorers() below.
Tuner ¶
One tuner class for every imputation model family - the same instance is meant to be reused across many Variables (e.g. via SRMI.Defaults), each getting its own independent search and its own result:
tuner = Tuner(
space=HyperparameterSpace(num_leaves=IntRange(2, 256), ...),
n_trials=50,
path_save_dir="tuner_outputs",
overwrite=True,
)
Parameters.LightGBM(tune=True, tuner=tuner, ...) # var A
Parameters.LightGBM(tune=True, tuner=tuner, ...) # var B, same tuner
Every routing call below (run/run_lightgbm/run_estimator) builds a FRESH optuna study - reusing one study across variables would silently mix each variable's (unrelated) trials into the same search, and study.best_params/best_value would then reflect whichever variable happened to score best overall, not the one you just tuned. Only space/n_trials/direction/seed/path_save_dir/overwrite are shared state; the study itself never is.
Three ways to use it, in increasing order of how much you write:
- run_lightgbm(train_data, test_data, base_params) - built-in LightGBM
routing (CV over lgb.basic.Dataset, scored via objective).
- run_estimator(estimator, X, y, cv=..., scoring=..., fit_params=...) -
built-in sklearn routing, argument names matching
sklearn.model_selection's own (GridSearchCV-style: clone(estimator)
.set_params(**trial_params) per trial).
- run(score_fn) - the escape hatch for anything else: score_fn(params:
dict) -> float is called once per trial with that trial's suggested
hyperparameters; fit/score however you need to, inside your own
closure. run_lightgbm/run_estimator are themselves just score_fn
closures built for you and passed to this same method.
run ¶
The generic primitive every routing method (including your own) is built on: score_fn(params: dict) -> float is called once per trial. Runs a FRESH study every call, saves the result to path_save (if set) and returns study.best_params.
run_estimator ¶
run_estimator(
estimator,
X,
y,
cv: int = 3,
scoring: str = "neg_root_mean_squared_error",
fit_params: dict | None = None,
direction: str = "maximize",
) -> dict
sklearn-style CV, argument names matching sklearn.model_selection's
own (GridSearchCV/cross_val_score: cv, scoring, fit_params) so
this reads like ordinary sklearn code. For each trial, clones
estimator (an unfitted, already-constructed sklearn-compatible
estimator - the same clone()+set_params() idiom GridSearchCV itself
uses), sets that trial's suggested hyperparameters on the clone,
fits/scores it across cv folds (a fixed fold assignment reused
across every trial, so score differences reflect hyperparameter
differences, not fold-split noise), and returns the mean fold score.
Splitting is done natively in polars: X (and y and any per-row
fit_params, e.g. sample_weight) are attached as columns on one
combined frame alongside a random fold-id column, then
partition_by(fold_id, include_key=False) both groups AND drops
the fold-id column in one step - so it can never leak into X as a
feature. This is this codebase's own native dataframe library
(used everywhere else in survey_kit) rather than reaching for
sklearn's private _safe_indexing (no backward-compatibility
guarantee per its own docstring) or supporting pandas (not used
anywhere in this codebase).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
estimator
|
sklearn-compatible estimator
|
An unfitted, already-constructed instance (e.g.
|
required |
X
|
DataFrame | ndarray
|
Already numeric/model-matrix-ready training data - a polars DataFrame (as Impute._prepare_tuning_data produces - needed to keep XGBoost's/CatBoost's native categorical dtype intact through to .fit()) or a plain numpy array (wrapped into a polars DataFrame internally). No pandas. |
required |
y
|
array - like
|
Target values, length matching X. |
required |
cv
|
int
|
Number of cross-validation folds, by default 3. |
3
|
scoring
|
str
|
Any sklearn scoring string (sklearn.metrics.get_scorer_names()), by default "neg_root_mean_squared_error". |
'neg_root_mean_squared_error'
|
fit_params
|
dict
|
Extra keyword arguments passed to every fold's .fit() call (e.g. {"sample_weight": w}) - split to each fold the same way X/y are, if array-like of matching length; passed as-is otherwise. By default None. |
None
|
direction
|
str
|
"maximize" or "minimize" - by default "maximize" (matching sklearn's own scorer convention, where higher is better - hence "neg_..." for error metrics). |
'maximize'
|
run_lightgbm ¶
run_lightgbm(
train_data,
test_data,
base_params: dict,
objective: Objective | None = None,
) -> dict
LightGBM's own CV: for each trial, train on train_data (an already-
built lgb.basic.Dataset) with base_params merged with that trial's
suggested hyperparameters, and score the fitted model's predictions
on test_data against test_data.label via objective (falls back to
this tuner's own self.objective, set at construction, if not given
here).
load_tuned_params ¶
Read back a hyperparameter dict previously saved by Tuner.run() (via path_save) - returns None if path is empty or nothing has been saved there yet (e.g. the very first run, before any tuning pass has completed). Used both by SRMI's own tune-before-run preprocessing (to check whether a fresh tuning pass is even needed) and by the actual per-iteration fit (to re-load the tuned values fresh from disk every time, rather than trying to carry a fitted-in-memory value across a deepcopy(variable) or a parallel worker process boundary).