Skip to content

Imputation/SRMI

SRMI

Bases: Serializable

Sequential Regression Multiple Imputation (SRMI) class for handling missing data imputation.

This class manages the complete SRMI process including variable setup, model configuration, parallel execution, and result management across multiple implicates and iterations.

Construction groups related settings into sub-objects - see init for the full parameter list:

df, variables, index, imputation_stats : flat, core settings replication : SRMI.Replication - n_implicates / n_iterations / seed parallel : SRMI.Parallel - enabled / variables_per_job / call_inputs / testing storage : SRMI.Storage - path_model / model_name / force_start / save_every_variable / save_every_iteration bootstrap : SRMI.Bootstrap - enabled / index / where defaults : SRMI.Defaults - weight / model / joint / selection / preselection / modeltype / parameters / ordered_categorical (fallback values applied to each Variable added via AddVariable() when that Variable doesn't specify its own)

Raises:

Type Description
Exception

If replication.n_implicates < 1 or replication.n_iterations < 1 If path equals path_model_new in load_to_continue_prior

Examples:

Basic usage:

>>> srmi = SRMI(
...     df=data,
...     variables=[var1, var2],
...     replication=SRMI.Replication(n_implicates=5, n_iterations=10),
...     storage=SRMI.Storage(path_model="/path/to/model"),
... )
>>> srmi.run()

With parallel execution:

>>> srmi = SRMI(
...     df=data,
...     variables=vars_list,
...     replication=SRMI.Replication(n_implicates=5, n_iterations=10),
...     parallel=SRMI.Parallel(enabled=True, call_inputs=CallInputs(n_cpu=4, mem_in_mb=5000)),
... )

simple_model classmethod

simple_model(
    df: IntoFrameT,
    index: list[str] | str | None = None,
    variables_to_impute: list[str] | None = None,
    classes: dict[str, Class] | None = None,
    auto_binary: bool = True,
    ordered_categories: dict[str, list] | None = None,
    model: dict[Class, ModelType | tuple] | None = None,
    yn_pairs: dict[str, str] | None = None,
    exclude: dict[str, list[str]] | None = None,
    exclude_global: list[str] | None = None,
    categorical_predictors: list[str] | None = None,
    group_levels: list[str] | str | None = None,
    replication: Replication = None,
    parallel: Parallel = None,
    storage: Storage = None,
    bootstrap: Bootstrap = None,
) -> SRMI

survey_kit's equivalent of mice's mice(data, m=5) one-liner on-ramp - point it at a dataframe and get back a ready-to-.run() SRMI with sensible per-variable defaults picked for you, that you can inspect/tweak on .variables before calling .run() yourself. This is a thin wrapper: all the actual defaulting logic (which columns need imputing, binary vs. continuous, ModelType/Parameters per Class, categorical-predictor handling, yn_pairs via Variable.two_part()) lives in utilities/auto_detect.py's build_simple_model_variables() - see its own docstring and this module's docstring for the full design rationale. Nothing here does anything Variable/ Parameters/SRMI couldn't already do by hand - for anything this doesn't cover, hand-build a Variable and pass it via SRMI(variables=[...]) directly instead.

Parameters:

Name Type Description Default
df IntoFrameT

The data to impute.

required
index list[str] | str | None

Same columns you'd pass to SRMI(index=...) directly (a row identifier is auto-generated if you don't give one, same as SRMI's own default) - also auto-excluded from every variable's predictor list here. By default None.

None
variables_to_impute list[str] | None

Exactly which columns to impute - if given, no auto- discovery happens beyond this list (nothing else is scanned for missingness). If None (the default), every column (other than index) with any missing values is imputed.

None
classes dict[str, Class] | None

{var_name: Variable.Class} - declares a variable's statistical type explicitly. Any imputed variable NOT listed here gets auto-classified as binary (boolean dtype, or numeric with only 0/1 values present) or continuous (auto_binary) - NEVER categorical; a variable that's really ordered_categorical or unordered_categorical always needs an explicit entry here, since category order (or "this numeric-looking column is actually a category code") can't be inferred from the data alone. By default None.

None
auto_binary bool

Whether an unclassified variable gets checked for binary-ness at all (see classes) - if False, every unclassified variable defaults straight to continuous. By default True.

True
ordered_categories dict[str, list] | None

{var_name: [ordered levels]} - required for any variable classes declares ordered_categorical; raises a clear error if missing. By default None.

None
model dict[Class, ModelType | tuple] | None

Per-Class model override. A value can be a bare Variable.ModelType (uses a built-in sensible default Parameters for it) or a (ModelType, parameters_dict) tuple (uses your parameters exactly). Unlisted classes use the built-in defaults: LightGBM for binary/continuous, RandomForest-backed (Multinomial/OrderedCategorical's own default estimator) for both categorical classes. By default None.

None
yn_pairs dict[str, str] | None

{value_var: yn_var} - route this variable through Variable.two_part() (semicontinuous/hurdle imputation) instead of a plain single Variable; yn_var must already exist as a column in df. Processed regardless of whether variables_to_impute would otherwise have included value_var/yn_var. By default None.

None
exclude dict[str, list[str]] | None

{impute_var: [vars]} - predictor columns to exclude for just that one variable. By default None.

None
exclude_global list[str] | None

Columns never used as a predictor for ANY variable built here. By default None.

None
categorical_predictors list[str] | None

Predictor columns that must be treated as categorical wherever they're used - passed as categorical_feature=... for a native-categorical modeltype (LightGBM/XGBoost/ CatBoost), or one-hot-encoded via a C(...) formula term otherwise (RandomForest/Multinomial/OrderedCategorical's default estimator/SklearnModel have no native categorical handling). By default None.

None
group_levels list[str] | str | None

Passed as group_levels=[...] to every built variable whose modeltype supports it (Regression/RandomForest/XGBoost/ CatBoost/SklearnModel - NOT the binary/continuous default of LightGBM, nor Multinomial/OrderedCategorical - a variable using one of those logs that group_levels was ignored for it). Also auto-folded into exclude_global, so it's never also used as an ordinary predictor. By default None.

None
replication Replication | Parallel | Storage | Bootstrap

Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults).

None
parallel Replication | Parallel | Storage | Bootstrap

Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults).

None
storage Replication | Parallel | Storage | Bootstrap

Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults).

None
bootstrap Replication | Parallel | Storage | Bootstrap

Same as passing these directly to SRMI(...) - unrelated to the auto-detection above. By default None (each group's own defaults).

None

Returns:

Type Description
SRMI

Constructed, not yet .run().

run

run() -> None

Execute the SRMI imputation process.

Orchestrates the complete imputation workflow including initialization, preprocessing, and execution in parallel or sequential mode.

Notes

The method performs these steps: 1. Creates folders and initializes implicates to be run 2. Preprocesses data (variable selection, hyperparameter tuning) 3. Runs imputation in parallel or sequential mode 4. Saves results and statistics

For parallel execution, creates job files for each iteration and variable subset. For sequential execution, runs each implicate directly.

convergence

convergence(
    diagnostic: str = "all", parameter: str = "mean"
) -> IntoFrameT

Convergence diagnostics for the imputed variables, across implicates and iterations - matches mice's own convergence() function (R/convergence.R) as closely as possible, including the exact algorithm it delegates to for the potential scale reduction factor (rstan::Rhat(), i.e. the rank-normalized, folded, split-Rhat of Vehtari et al. 2021 - see utilities/convergence_diagnostics.py's module docstring for the verified-against-source details).

Callable any time after at least 3 iterations of at least 2 implicates have run (mid-run is fine, not just after a completed SRMI) - it reads whatever each Impute.chain_mean/ chain_std has already accumulated on self.implicates, the same mean/std of each variable's own newly-imputed values every iteration that Impute._post_impute_statistics computes (and that already respects the Variable's own Where restriction, since it's read from df_impute, which Impute.df_impute() has already filtered by Where before this ever sees it).

Parameters:

Name Type Description Default
diagnostic str

"all" (both ac and psrf), "ac" (lag-1 autocorrelation only), or "psrf"/"gr" (potential scale reduction factor only). By default "all".

'all'
parameter str

"mean" or "sd" - diagnose the chain means or the chain standard deviations. By default "mean".

'mean'

Returns:

Type Description
IntoFrameT

One row per (iteration, variable) - columns ".it", "vrb", and "ac"/"psrf" per diagnostic, same names mice uses. NaN wherever a variable has no numeric/boolean imputed values to track for that (iteration, implicate) - e.g. a Multinomial()/ OrderedCategorical() category-label target, or an iteration where that variable had nothing to impute.

plot_convergence

plot_convergence(
    parameter: str = "both", path: str | None = None
) -> "plotly.graph_objects.Figure"

Plot the trace lines of the SRMI algorithm - matches mice's own plot(imp) (plot.mids(), R/mids.R) as closely as possible: for each imputed variable, one line per implicate, the chain mean (and/or chain sd) against iteration number. On convergence, the lines within a panel should intermingle and be free of any trend - the same "worm plot" reading as mice's own.

Purely a plotting convenience on top of the same data convergence() reads (Impute.chain_mean/chain_std, accumulated on self.implicates) - there's no "auto-plot" flag anywhere in SRMI's own run() - call this whenever you want, including on an SRMI you've just SRMI.load()ed in a completely different process/environment from the one that ran the imputation (e.g. one where plotly wasn't installed at run time, but is now).

Requires the optional 'plotly' package - not a survey_kit dependency at all (imputation itself never needs it), so nothing about running SRMI requires having it installed; only calling this specific method does.

Parameters:

Name Type Description Default
parameter str

"both" (mice's own default - one row of panels for chain means, one for chain sds), "mean", or "sd". By default "both".

'both'
path str | None

If given, also save the figure there as a self-contained HTML file (fig.write_html) - no extra dependency beyond plotly itself needed for that, unlike a static image export (which would need kaleido too). By default None (don't save - just return the Figure; display it yourself, e.g. fig.show() in a script or automatically in a notebook).

None

Returns:

Type Description
Figure

Always returned (even when path is also given) so you can further customize it, .show() it, or save it yourself in a different format.

plot_imputation_quality

plot_imputation_quality(
    variable: str | int | list[str | int] | None = None,
    kind: str = "density",
    sample_k: int | None = None,
    seed: int | None = None,
    path: str | None = None,
) -> "plotly.graph_objects.Figure"

Plot observed vs. imputed values - a DIFFERENT question from plot_convergence()'s "did the chain stabilize": here it's "do the imputed values look plausible next to the observed ones." Matches mice's own densityplot()/stripplot()/bwplot() as closely as possible (see utilities/quality_diagnostics.py's module docstring for the verified-against-source convention): one group for the truly observed values (pooled once), plus one group per implicate holding ONLY that implicate's own newly imputed values - never the whole column.

Not built here (a materially bigger lift - needs a fitted propensity/detrending model, not just a reshape): mice's propensity-score xyplot() and its own "worm plot" (a detrended Q-Q plot conditional on a covariate - unrelated to the trace lines plot_convergence() draws, despite the similar-sounding name).

No "auto-plot" flag in run() here either, same as plot_convergence() - call this whenever you want, including on an SRMI you've just SRMI.load()ed.

Requires the optional 'plotly' package, same as plot_convergence() - raises a clear ImportError if it isn't installed, only when this is actually called.

Parameters:

Name Type Description Default
variable str | int | list[str | int] | None

Which variable(s) to plot - a name (impute_var), a 0-indexed position in self.variables, a list mixing either, or None for every variable in self.variables. By default None (all).

None
kind str

"density" (kernel density per group - numeric/boolean variables only, silently skips a group with fewer than 2 distinct finite values, same wall mice's own densityplot() hits with no workaround), "strip" (every individual point, one column per group), or "box" (five-number-summary box plot per group). By default "density".

'density'
sample_k int | None

Only meaningful for kind="strip" - randomly sample at most this many points per (group, variable) to avoid overplotting a large dataset (mice's own guidance: stripplot is best for small datasets, use bwplot/box for large ones - this is the other way to cope with a large one and still see individual points). Ignored for "density"/"box", which should always use every point. By default None (no sampling).

None
seed int | None

Seed for the sample_k random sample. By default None.

None
path str | None

If given, also save the figure there as a self-contained HTML file. By default None.

None

Returns:

Type Description
Figure

plot_propensity

plot_propensity(
    variable: str | int | list[str | int] | None = None,
    kind: str = "density",
    n_bins: int = 10,
    cv_folds: int = 5,
    lightgbm_parameters: dict | None = None,
    residual_model_parameters: dict | None = None,
    path: str | None = None,
) -> "plotly.graph_objects.Figure"

Response-propensity diagnostic - a different, conditional question from plot_imputation_quality()'s marginal density/ strip/box comparison: instead of "does the imputed marginal distribution look like the observed marginal distribution" (which a correct MAR imputation can legitimately fail, since missing rows can differ systematically on the predictors), this asks "conditional on how similar a row's covariate profile is to a typically-missing row (its response propensity, from a binary LightGBM model of the imputation_flag on that variable's own predictors), does the imputed value track the observed one." See utilities/propensity_diagnostics.py's module docstring for the full rationale, including how kind="density" matches the diagnostic used in Raghunathan & Bondarenko (2007)/Bondarenko & Raghunathan (2016) - and in the user's own SRMI/CPS-ASEC paper (Hokayem, Raghunathan & Rothbaum) - rather than an invented convenience.

No "auto-plot" flag, same as plot_convergence()/ plot_imputation_quality() - call this whenever you want, including on an SRMI you've just SRMI.load()ed. Requires the optional 'plotly' package, same as the other two.

Parameters:

Name Type Description Default
variable str | int | list[str | int] | None

Which variable(s) to plot - a name (impute_var), a 0-indexed position in self.variables, a list mixing either, or None for every variable in self.variables. By default None (all).

None
kind str

"density" (default) - regress y on the propensity (a LightGBM regression, one continuous feature, fit on the observed rows only), then compare the KERNEL DENSITY of the residuals - observed vs. each implicate's own imputed rows, scored against that same fit. Similar location AND spread across groups is the good outcome; a shifted or differently-spread residual density for an implicate suggests the imputation model is missing something the missingness mechanism itself depends on. The observed group's own residuals come from cv_folds-way cross- validated (out-of-fold) predictions specifically to avoid biasing them tighter than the (always out-of-sample) implicate residuals - see propensity_diagnostics.propensity_residual_data()'s docstring for why that matters for a single-feature model. "binned_mean" - a simpler, cruder alternative: bin rows by propensity (quantiles of the pooled sample) and compare mean(y), not residuals, within each bin.

'density'
n_bins int

Only used by kind="binned_mean" - number of quantile bins of the pooled propensity to group rows into, by default 10. Must be >= 2.

10
cv_folds int

Only used by kind="density" - number of cross-validation folds for the observed group's out-of-fold residuals, by default 5. Must be >= 2 to matter; a variable with too few observed rows for the requested fold count (fewer than 2*cv_folds) is silently skipped (logged), same as any other insufficient-data case elsewhere in these diagnostics.

5
lightgbm_parameters dict | None

Overrides for the propensity model's LightGBM parameters - merged onto propensity_diagnostics.DEFAULT_PROPENSITY_PARAMETERS (a deliberately shallow/regularized default - an unregularized GBM overfits the in-sample propensity toward 0/1 and collapses the diagnostic). By default None. Each variable's own categorical_feature (from its CatBoost()/ XGBoost()/RandomForest()/SklearnModel()/LightGBM() parameters, whichever it uses) is detected and passed to the propensity model automatically - no need to repeat it here unless you want to override it.

None
residual_model_parameters dict | None

Only used by kind="density" - overrides for the y~propensity LightGBM regression's own parameters, merged onto propensity_diagnostics.DEFAULT_RESIDUAL_MODEL_PARAMETERS. By default None.

None
path str | None

If given, also save the figure there as a self-contained HTML file. By default None.

None

Returns:

Type Description
Figure

Variable

Bases: Serializable

Class

Bases: Enum

A variable's statistical TYPE (separate from ModelType, which picks - a specific fitting algorithm for that type). Used by SRMI.simple_model()/utilities/auto_detect.py to pick a sensible default ModelType per variable: binary/continuous can be told apart automatically (dtype/0-1-only check); ordered_categorical and unordered_categorical can never be inferred from the data alone (there's no way to know intended category order, or that a numeric-looking column is really a category code, without being told) - a variable of either categorical kind always needs an explicit Class declaration.

PrePost

Namespace within Variable class for handling pre and post imputation operations.

Currently that can be: 1) a Narwhals Expr (NarwhalsExpression) which is anything you can put in nw.from_native(df).with_columns() 2) a python function handle and parameters which allows you to call an arbitrary function

Function

NarwhalsExpression

Bases: Serializable

PolarsExpression

Bases: Serializable

Predictors

Bases: Serializable

Which predictor variables enter (or are forced into/out of) this variable's model.

Sample

Bases: Serializable

Which rows this variable's imputation applies to.

Transforms

Bases: Serializable

Operations to run before/after this variable's imputation each iteration (pre/post), plus once-only bookends around the whole implicate (pre_initialize before the first iteration, post_finalize after the last).

from_legacy classmethod

from_legacy(
    impute_var: str = "",
    Where: Expr | None = None,
    Where_impute: Expr | None = None,
    Where_predict: Expr | None = None,
    Where_predict_only_when_not_imputed: bool = False,
    bimpute_if_missing: bool = True,
    preFunctions=None,
    postFunctions=None,
    preFunctions_initialize_implicate=None,
    predictors_exclude: list = None,
    predictors_exclude_first_iteration: list = None,
    predictors_require: list = None,
    weight: str = "",
    joint: dict = None,
    header: str = "",
    model: str = "",
    selection: Selection = None,
    preselection: Selection = None,
    modeltype: ModelType = None,
    modelfunction=None,
    parameters: dict = None,
    By: list = None,
) -> Variable

Construct a Variable from a flat set of keyword arguments, rather than the grouped sample=Variable.Sample(...)/ transforms=Variable.Transforms(...)/ predictors=Variable.Predictors(...) form Variable() itself takes.

two_part classmethod

two_part(
    df: IntoFrameT,
    impute_var: str,
    model: list[str] | str,
    modeltype: ModelType = ModelType.Regression,
    parameters: dict | None = None,
    yn_model: list[str] | str | None = None,
    yn_modeltype: ModelType | None = None,
    yn_parameters: dict | None = None,
    yn_var: str | None = None,
    yn_missing: Expr | None = None,
    value_if_no: float | None = 0,
    weight: str = "",
    By: list[str] | str | None = None,
) -> tuple[IntoFrameT, list[Variable]]

Shortcut for semicontinuous (point mass at zero + continuous) two-part imputation: a binary y/n ("is impute_var nonzero") variable, imputed first, then impute_var itself, restricted to the y/n==True population - the standard two-part/hurdle approach (as opposed to just PMM/leaf-matching the raw variable directly, which reproduces the point mass for free but assumes a single model/predictor set explains both the participation and intensity margins - see the two_part design discussion this wraps up).

Takes df (needed to derive the y/n column - see below) and returns (df, variables): df with the y/n column added, and a plain list, always safe to variables.extend(...) or loop over - normally [yn_variable, value_variable] (order matters: yn must precede value in your SRMI variables list, so value sees this iteration's fresh yn draw, not last iteration's), but sometimes shorter - see yn_var and the HotDeck/StatMatch note below. Use the RETURNED df (not your original) to construct SRMI - SRMI.init needs every impute_var to already exist as a real column up front, even the y/n one, so it can't be created lazily via a hook the way the per-iteration consistency fixups below are.

Mechanics (mostly via existing Variable.Transforms hooks, no per-call special-casing in impute.py): - If creating y/n (yn_var not given): the column is derived once, right here, straight into the returned df (not via pre_initialize - the derivation is a fixed function of df's own observed values, identical for every implicate, so there's nothing to recompute per-implicate): null wherever impute_var is null (or, if yn_missing is given, ALSO wherever that expression is true - additive, never narrower, since a null impute_var can never tell us whether y/n was really True or False), else impute_var != 0. - If yn_var IS given but still has missing values of its own: value_variable's Where (below) would otherwise silently exclude those rows from ever being touched at all (a null never satisfies a boolean filter) - so a real yn_variable still gets built for it, using the same modeltype/ parameters derivation as the create-our-own case, just reading/imputing the caller's own column in place rather than deriving it, and never dropping it at the end (not ours to clean up). If yn_var has no missing values at all, none of this applies - value_variable alone is returned. - Whenever a real yn_variable is built (either case above) and its own donation is genuine (pmm/leaf, or HotDeck/ StatMatch - always donation-based), impute_var itself rides along in its donate_list - so a row whose y/n was just resolved gets its value from that SAME matched donor, rather than from a possibly different one value_variable's own later pass would find. A harmless no-op when yn's error draw is Random instead, which never donates anything. - value_variable.sample.Where restricts value's donor pool AND recipients to yn==True rows. - value_variable.transforms.pre runs every iteration, before value's own fit/donation: wherever yn is currently True but value == 0 (stale from a prior iteration when yn was False), null it out (so it's a genuine recipient again this iteration); wherever yn is currently False, force value to value_if_no. - yn_variable.transforms.post_finalize drops the scratch y/n column once, after the implicate's last iteration - only when this function created it (see yn_var).

Parameters:

Name Type Description Default
df IntoFrameT

Source data - read to derive the y/n column, and to return the augmented copy you should actually build SRMI from.

required
impute_var str

The semicontinuous target.

required
model list[str] | str

Predictors for impute_var (value). Also yn's predictors, unless yn_model overrides them.

required
modeltype ModelType

value's modeltype. Supported: Regression, pmm, LightGBM (fit a real Logit/binary-objective variant for yn); RandomForest, XGBoost, CatBoost, SklearnModel (yn reuses the SAME estimator factory via ModelType.OrderedCategorical with categories=[False, True] - no regressor/classifier estimator swap, which isn't generically possible: sklearn/ XGBoost/CatBoost each have model-family-specific hyperparameters, like criterion/objective/loss_function, that don't transfer across that boundary, and there's no way to do it at all for SklearnModel's arbitrary factory); HotDeck, StatMatch (donation doesn't care about target dtype, no swap needed - see the collapse behavior below). Anything else (Multinomial, OrderedCategorical) raises - not semicontinuous-shaped as value's own type. By default Variable.ModelType.Regression.

Regression
parameters dict

value's parameters (e.g. Parameters.RandomForest(...)). By default None ({}).

None
yn_model list[str] | str | None

Predictors for yn, if they should differ from model - the participation and intensity margins often don't share predictors/mechanism. By default None (same as model).

None
yn_modeltype ModelType | None

Explicit override for yn's modeltype - pass together with yn_parameters for full manual control. By default None (derived from modeltype per the modeltype docstring above).

None
yn_parameters dict | None

Explicit override for yn's parameters. If given without yn_modeltype, yn_modeltype defaults to modeltype (or OrderedCategorical, if modeltype is one of the four that reuse it) rather than being derived further - if you're supplying parameters yourself, supply the modeltype too if it's not that default. By default None (derived).

None
yn_var str | None

Reuse an existing y/n column instead of deriving one from impute_var. If it's already fully observed, this function builds ONLY value_variable (a 1-item list), restricted to yn_var==True. If it still has missing values of its own, this function ALSO builds a yn_variable to resolve them (using the same modeltype/yn_modeltype/yn_parameters derivation as the create-our-own case - see the mechanics note above), so the returned list is still [yn_variable, value_variable] in that case - it's only "your own concern" when there's genuinely nothing left for it to do. Never dropped at the end either way - it's your column. By default None (create one, named f"{imputevar}_yn_").

None
yn_missing Expr | None

Only meaningful when yn_var is not given. Rows where this is true are ALSO treated as yn-missing, on top of impute_var.is_null() (additive, not a replacement - see the mechanics note above for why). By default None.

None
value_if_no float | None

What value becomes on yn==False rows, every iteration. By default 0 (pass None to leave it null instead).

0
weight str

Applied identically to both variables. By default "".

''
By list[str] | str | None

Applied identically to both variables. By default None.

None

Returns:

Type Description
tuple[IntoFrameT, list[Variable]]

(df, variables) - df augmented with the y/n column (only when this function created one - unchanged otherwise); variables is [yn_variable, value_variable] normally, [value_variable] alone only when either (a) yn_var was given and is already fully observed (nothing left for a yn_variable to do), or (b) modeltype is HotDeck/StatMatch and there's no signal at all (no yn_var, yn_missing, yn_model, yn_modeltype, or yn_parameters) that the two margins should be modeled differently - a single donation pass on value, unrestricted, already reproduces the point mass for free in that case, so a redundant yn model is skipped (with a warning).

union_required_predictors

union_required_predictors(formula: str) -> str

Add any predictors_require variables not already present in formula.

Used after variable selection (LASSO/stepwise) replaces the model formula, to guarantee required predictors survive selection - the selection methods themselves never see predictors_require (they don't receive the Variable object), so this has to happen here, in the caller, once selection has returned.

Parameters:

Name Type Description Default
formula str

A model formula, typically the one returned by Selection.run()/lasso()/stepwise().

required

Returns:

Type Description
str

formula, with any missing required predictors added.

Parameters

Factory class for creating parameter dictionaries for different imputation methods.

Provides static methods to generate properly formatted parameter dictionaries with validation and default values for each imputation approach.

ErrorDraw

Bases: Enum

Random = 0 pmm = 1 leaf = 2

RegressionModel

Bases: Enum

OLS = 0

Probit = 1

Logit = 2

CatBoost staticmethod

CatBoost(
    parameters: dict | None = None,
    error: ErrorDraw = ErrorDraw.pmm,
    random_share: float = 1.0,
    cv_folds: int = 0,
    categorical_feature: list[str] | str | None = None,
    group_levels: list[str] | None = None,
    group_shrinkage_k: float = 10.0,
    parameters_pmm: dict = None,
    tune: bool = False,
    tuner=None,
) -> dict

Parameters for CatBoost-based imputation (mean regression only).

Parameters:

Name Type Description Default
parameters dict

Keyword arguments passed straight through to catboost.CatBoostRegressor(**parameters), by default None (CatBoostRegressor's own defaults, plus verbose=0).

None
error ErrorDraw

Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring.

pmm
random_share float

Fraction of data to use for fitting, by default 1.0.

1.0
cv_folds int

See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off).

0
categorical_feature list[str] | str | None

Predictor column names to treat as native categoricals (CatBoost's own ordered-target-statistic categorical handling, not one-hot encoding) - see XGBoost()'s categorical_feature docstring for why this is separate from a formula's C(...) syntax. By default None (no native categoricals).

None
group_levels list[str] | None

See Regression()'s group_levels docstring, by default None (off).

None
group_shrinkage_k float

See Regression()'s group_shrinkage_k docstring, by default 10.0.

10.0
parameters_pmm dict

PMM parameters if using PMM error drawing, by default None.

None
tune bool

Whether to tune hyperparameters for this run, by default False. No effect if tuner is None. See utilities.tuning.Tuner - SRMI runs a Tuner.run_estimator() pass (over tuner's own HyperparameterSpace, against the constructed CatBoostRegressor as the template estimator) before the SRMI run starts.

False
tuner Tuner

Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None.

None

Returns:

Type Description
dict

CatBoost parameter dictionary

HotDeck staticmethod

HotDeck(
    model_list: list[str] | list[list[str]] = None,
    donate_list: list | None = None,
    n_hotdeck_array: int = 3,
    sequential_drop: bool = True,
) -> dict

Parameters for hot HotDeck imputation

Parameters:

Name Type Description Default
model_list list[str] | list[list[str]]

Each model is a list of variables that are used as match keys. model_list can either be a list of strings (the model itself) or it can be a list of lists of strings (sequential hot deck to match on)

None
donate_list list

Additional variables to impute together, by default None I.e., you can predict earnings amount to find a donor, but then also impute hours worked and weeks worked with it.

None
n_hotdeck_array int

Size of hot deck donor arrays, by default 3

3
sequential_drop bool

Drop variables sequentially until matches found, by default True If model_list is a list of strings (one model), should we sequentially drop the last variable until all recipients find a donor? Makes it easier to set the hot deck/stat match up. Whatever recipients are STILL unmatched after every model in model_list (including every dropped-down level of this cascade) has been tried get matched fully at random instead, with a logged warning, rather than being left unmatched forever - see impute.py's hotdeck()/statmatch(), the guaranteed-to-succeed last resort applies regardless of sequential_drop or what model_list contains.

True

Returns:

Type Description
dict

Hot deck parameter dictionary

LightGBM staticmethod

LightGBM(
    tune: bool = False,
    parameters: dict | None = None,
    tuner=None,
    quantiles: list = None,
    error: ErrorDraw = ErrorDraw.pmm,
    cv_folds: int = 0,
    parameters_pmm: dict | None = None,
) -> dict

Parameters for LightGBM-based imputation.

Parameters:

Name Type Description Default
tune bool

Whether to tune hyperparameters for this run, by default False. No effect if tuner is None.

False
parameters dict | None

LightGBM model parameters, by default None

None
tuner Tuner

Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes - see utilities.tuning.Tuner / HyperparameterSpace), by default None

None
quantiles list

Quantiles for quantile regression, by default None

None
error ErrorDraw

How to convert yhat from LGBM into imputes. If pmm, draw from nearest yhat neighbors, for example. The default is ErrorDraw.pmm.

pmm
cv_folds int

See Parameters._tabular_ml_params's cv_folds docstring - same tabular-ML cv_folds mechanism RandomForest()/XGBoost()/ CatBoost()/SklearnModel() use, by default 0 (off).

0
parameters_pmm dict | None

PMM parameters if using PMM error drawing, by default None. "winsor" is dropped even if present - unlike pmm()/Regression(), LightGBM's own fit path never winsorizes impute_var before training (that's only done by _pmm_winsorize_for_fit, called from pmm()/regression() - the latter also covers RandomForest/ XGBoost/CatBoost/SklearnModel, which route through regression(), but LightGBM has its own separate fit path that doesn't).

None

Returns:

Type Description
dict

LightGBM parameter dictionary

Raises:

Type Description
Exception

If quantiles are not between 0 and 1

Multinomial staticmethod

Multinomial(
    parameters: dict | None = None,
    donate_list: list[str] | None = None,
    donate_by: list[str] | str | None = None,
    random_share: float = 1.0,
) -> dict

Parameters for imputing an unordered categorical variable with 3+ levels via a RandomForestClassifier and donor matching on leaf co-occurrence - see imputation/utilities/leaf_donor_matching.leaf_cooccurrence_match for the donor-selection mechanism (the same one mice's rf method uses: pool donors sharing a leaf with the recipient across every tree, draw one uniformly at random).

This is a distinct imputation shape from RandomForest()/XGBoost()/CatBoost()/SklearnModel() above - those are all mean regression (a single continuous yhat, optionally PMM-matched); this is genuinely multi-class classification, with the donor match itself driven by which trees group two rows together, not by distance on a predicted scalar. There's no error=/cv_folds= here for that reason - PMM-style matching on a scalar prediction and "refit on folds to correct in-sample bias" both assume a scalar yhat, which doesn't exist for this method.

Parameters:

Name Type Description Default
parameters dict

Keyword arguments passed straight through to sklearn.ensemble.RandomForestClassifier(**parameters), by default None (RandomForestClassifier's own defaults). No native categorical predictor handling (same restriction as RandomForest()) - model= can be a plain column list or an R-style formula string; either way, a categorical predictor needs to end up numeric before reaching this model, via the formula's own C(...) one-hot encoding or, for the list form, by already being numeric-coded.

None
donate_list list[str]

Additional variables to impute together from the same matched donor, by default None.

None
donate_by list[str] | str | None

Grouping variable(s) - donors are only matched within the recipient's own group, by default None.

None
random_share float

Fraction of data to use for fitting, by default 1.0.

1.0

Returns:

Type Description
dict

Multinomial parameter dictionary

NearestNeighbor staticmethod

NearestNeighbor(
    match_to: str | list[str],
    logit: bool = False,
    parameters_pmm: dict = None,
) -> dict

Convenience builder for nearest-neighbor PMM matching, purely for readability/discoverability - there's no dedicated Variable.ModelType.NearestNeighbor; this just returns Parameters.Regression()'s own dict shape, so use it with Variable.ModelType.Regression. Matching on a fitted OLS/Logit prediction (a principled notion of "similar") is strictly better than matching on raw, unweighted, unlearned multivariate distance across match_to's predictors directly, so this always routes through Regression rather than a distance-based implementation of its own.

Parameters:

Name Type Description Default
match_to str | list[str]

The predictor(s) to match on. NOT stored anywhere in the returned dict - Parameters.XXX() functions never set Variable.model themselves (true of every one of them, not specific to this one), so pass this SAME list as the Variable's own model= too.

required
logit bool

impute_var is binary - fit Logit instead of OLS (still matched via PMM on the fitted probability, same as OLS's fitted mean). By default False (OLS).

False
parameters_pmm dict

knearest/donate_list/donate_by/winsor - same as pmm()/ Regression(). By default None (Parameters.pmm()'s own defaults).

None

Returns:

Type Description
dict

Parameters.Regression(model=OLS or Logit, error=pmm, ...)'s parameter dictionary.

OrderedCategorical staticmethod

OrderedCategorical(
    categories: list,
    parameters: dict | None = None,
    estimator: Callable[[], object] | None = None,
    error: ErrorDraw = ErrorDraw.pmm,
    random_share: float = 1.0,
    categorical_feature: list[str] | str | None = None,
    estimator_prepare_data: Callable[
        [object, object], tuple[object, object]
    ]
    | None = None,
    donate_list: list[str] | None = None,
    donate_by: list[str] | str | None = None,
    knearest: int = 10,
) -> dict

Parameters for imputing an ORDERED categorical variable (e.g. an education level or a Likert scale) - unlike Multinomial() (unordered, classification), this fits a mean-regression estimator against an integer rank encoding of categories (lowest/coarsest to highest/finest), then donates the REAL observed category from a matched donor - never the numeric rank, and never a category that wasn't actually observed, the same guarantee Multinomial() and PMM/leaf donor matching already give elsewhere.

This reuses the same donor-matching machinery RandomForest()/ XGBoost()/CatBoost() etc. use for their own error=pmm/leaf, just matched on the predicted rank instead of impute_var's own (non-numeric) value - see impute.py's ordered_categorical().

Parameters:

Name Type Description Default
categories list

Every observed value of impute_var, ordered from lowest/coarsest to highest/finest (e.g. ["less_than_hs", "hs_grad", "some_college", "college_grad"]). Imputation raises if any observed value isn't in this list.

required
parameters dict

Keyword arguments for the default estimator (sklearn.ensemble.RandomForestRegressor(**parameters)) - ignored if estimator is set. By default None.

None
estimator Callable[[], object]

Zero-arg factory for a fresh, unfitted mean-regression estimator (e.g. lambda: XGBRegressor(...) or lambda: CatBoostRegressor(...)), same shape as SklearnModel()'s factory. By default None (uses RandomForestRegressor(**parameters)).

None
error ErrorDraw

pmm (knearest on the predicted rank) or leaf (tree leaf co-occurrence - needs estimator to expose leaf indices, e.g. the default RandomForestRegressor, or your own XGBRegressor/CatBoostRegressor factory - see ErrorDraw.leaf's docstring). ErrorDraw.Random isn't meaningful here (there's no natural way to add continuous noise to a category and get a valid category back). By default ErrorDraw.pmm.

pmm
random_share float

Fraction of data to use for fitting, by default 1.0.

1.0
categorical_feature list[str] | str | None

Predictor columns to cast to a fixed-category dtype before fitting, for an estimator with native categorical handling (XGBoost/CatBoost) - see RandomForest()/XGBoost()'s docstring. By default None.

None
estimator_prepare_data Callable[[df_model, df_impute], (df_model, df_impute)]

See SklearnModel()'s prepare_data docstring. By default None.

None
donate_list list[str]

Additional variables to impute together from the same matched donor, by default None.

None
donate_by list[str] | str | None

Grouping variable(s) - donors are only matched within the recipient's own group, by default None.

None
knearest int

Only used when error=pmm - number of nearest neighbors (on predicted rank) to draw from, by default 10.

10

Returns:

Type Description
dict

OrderedCategorical parameter dictionary

RandomForest staticmethod

RandomForest(
    parameters: dict | None = None,
    error: ErrorDraw = ErrorDraw.pmm,
    random_share: float = 1.0,
    cv_folds: int = 0,
    group_levels: list[str] | None = None,
    group_shrinkage_k: float = 10.0,
    parameters_pmm: dict = None,
    tune: bool = False,
    tuner=None,
) -> dict

Parameters for RandomForestRegressor-based imputation (mean regression only - scikit-learn's RandomForestRegressor has no native quantile-loss support). No native categorical handling either - categorical predictors need to be encoded (e.g. via a "~...+C(var)+..." formula) before reaching this model, same as OLS/Logit.

Parameters:

Name Type Description Default
parameters dict

Keyword arguments passed straight through to sklearn.ensemble.RandomForestRegressor(**parameters), by default None (RandomForestRegressor's own defaults).

None
error ErrorDraw

Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring.

pmm
random_share float

Fraction of data to use for fitting, by default 1.0.

1.0
cv_folds int

See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off).

0
group_levels list[str] | None

See Regression()'s group_levels docstring, by default None (off).

None
group_shrinkage_k float

See Regression()'s group_shrinkage_k docstring, by default 10.0.

10.0
parameters_pmm dict

PMM parameters if using PMM error drawing, by default None.

None
tune bool

Whether to tune hyperparameters for this run, by default False. No effect if tuner is None. See utilities.tuning.Tuner - SRMI runs a Tuner.run_estimator() pass (over tuner's own HyperparameterSpace, against RandomForestRegressor(**parameters) as the template estimator) before the SRMI run starts.

False
tuner Tuner

Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None.

None

Returns:

Type Description
dict

RandomForest parameter dictionary

Regression staticmethod

Regression(
    model: RegressionModel = RegressionModel.OLS,
    error: ErrorDraw = ErrorDraw.pmm,
    random_share: float = 1.0,
    group_levels: list[str] | None = None,
    group_shrinkage_k: float = 10.0,
    parameters_pmm: dict = None,
) -> dict

Parameters for regression-based imputation.

Parameters:

Name Type Description Default
model RegressionModel

Type of regression model, by default RegressionModel.OLS

OLS
error ErrorDraw

Method for drawing errors, by default ErrorDraw.Random. ErrorDraw.leaf isn't usable here (OLS/Logit have no tree structure to match donors on) - see RandomForest()/ XGBoost()/CatBoost()/SklearnModel() for that. If pmm, draw from nearest yhat neighbors, for example. If Random, draw from observed errors for modeled observations

pmm
random_share float

Fraction of data to use for regression, by default 1.0 Use less memory by running the regression on a random subset?

1.0
group_levels list[str] | None

Nested grouping columns, COARSEST to FINEST (e.g. ["state", "county", "hhid"]), for a cheap shrinkage-heuristic stand-in for a random-intercept term - not a real mixed model, see Impute._nested_group_shrinkage's docstring for the exact mechanism and its limits. Re-estimated fresh from each SRMI iteration's fit residuals and persisted into the working data (as a plain column, the same way donate_list values ride along) for the next iteration to build on - no separate inner convergence loop. By default None (off).

None
group_shrinkage_k float

Shrinkage constant for group_levels - a group needs roughly this many observations before its own mean starts to dominate over being pulled toward 0. Only meaningful when group_levels is set. By default 10.0.

10.0
parameters_pmm dict

PMM parameters if using PMM error drawing, by default None

None

Returns:

Type Description
dict

Regression parameter dictionary

SklearnModel staticmethod

SklearnModel(
    factory: Callable[[], object],
    error: ErrorDraw = ErrorDraw.pmm,
    random_share: float = 1.0,
    cv_folds: int = 0,
    categorical_feature: list[str] | str | None = None,
    prepare_data: Callable[
        [object, object], tuple[object, object]
    ]
    | None = None,
    group_levels: list[str] | None = None,
    group_shrinkage_k: float = 10.0,
    parameters_pmm: dict = None,
    tune: bool = False,
    tuner=None,
) -> dict

Parameters for imputation using any sklearn-compatible estimator you bring yourself (mean regression only) - the escape hatch for anything without a dedicated RandomForest()/XGBoost()/CatBoost() function.

Parameters:

Name Type Description Default
factory Callable[[], object]

Zero-arg callable returning a fresh, unfitted sklearn-compatible estimator (supporting .fit(X, y) and .predict(X), plus .fit(..., sample_weight=...) if you're using a weight variable). E.g. lambda: MyModel(some_hyperparam=5).

required
error ErrorDraw

Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring.

pmm
random_share float

Fraction of data to use for fitting, by default 1.0.

1.0
cv_folds int

See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off).

0
categorical_feature list[str] | str | None

Predictor column names to cast to a fixed-category dtype (polars Enum) in the model matrix before fitting - useful if your own estimator wants that dtype for native categorical handling, the same way XGBoost/CatBoost do. This only casts the dtype; if your model needs the categorical columns communicated some other way too (e.g. a constructor kwarg naming them), use prepare_data instead/as well. By default None.

None
prepare_data Callable[[df_model, df_impute], (df_model, df_impute)]

Full control over preparing the model matrix right before fitting/predicting - called once, on both the donor pool and recipient frames together (so anything derived from both, like fixed Enum categories, stays consistent between them), after the regular formula/model-matrix construction and before your factory's model is fit. By default None (no extra prep beyond categorical_feature, if any).

None
group_levels list[str] | None

See Regression()'s group_levels docstring, by default None (off).

None
group_shrinkage_k float

See Regression()'s group_shrinkage_k docstring, by default 10.0.

10.0
parameters_pmm dict

PMM parameters if using PMM error drawing, by default None.

None
tune bool

Whether to tune hyperparameters for this run, by default False. No effect if tuner is None. See utilities.tuning.Tuner - SRMI runs a Tuner.run_estimator() pass (over tuner's own HyperparameterSpace, against factory() as the template estimator) before the SRMI run starts. Your estimator needs a real sklearn .set_params()/.get_params() (true of anything built on sklearn.base.BaseEstimator) for the tuned hyperparameters to actually apply.

False
tuner Tuner

Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None.

None

Returns:

Type Description
dict

Custom sklearn-model parameter dictionary

StatMatch staticmethod

StatMatch(
    model_list: list = None,
    donate_list: list = None,
    sequential_drop: bool = False,
)

Parameters for hot statistical match imputation

Stat match and hot deck are basically the same, but the hot deck iterates over the data carrying arrays of possible donor values whereas the stat match just does a join of donors and recipients

Parameters:

Name Type Description Default
model_list list[str] | list[list[str]]

Each model is a list of variables that are used as match keys. model_list can either be a list of strings (the model itself) or it can be a list of lists of strings (sequential hot deck to match on)

None
donate_list list

Additional variables to impute together, by default None I.e., you can predict earnings amount to find a donor, but then also impute hours worked and weeks worked with it.

None
sequential_drop bool

Drop variables sequentially until matches found, by default False If model_list is a list of strings (one model), should we sequentially drop the last variable until all recipients find a donor? Makes it easier to set the hot deck/stat match up.

False

Returns:

Type Description
dict

stat match parameter dictionary

TwoSampleRegression staticmethod

TwoSampleRegression(
    model: "Parameters.RegressionModel" = None,
    is_boolean: bool = False,
    bins: int = 10,
    bin_by: list[str] | str | None = None,
    percentile_cuts: list[float] | None = None,
    save_percentile_cuts: bool = False,
    round_impute_var_digits: int = 4,
    continuous_qtiles_y_cuts: list[float] | None = None,
    continuous_qtiles_interpolate_by_bin: bool = False,
    min_n_x_var: int = 0,
    draw_error: bool = False,
    random_share: float = 1.0,
    save_disclosure_support: bool = False,
    path_save: str = "",
    path_load: str = "",
    load_from_save: bool = False,
) -> dict

Parameters for two-sample regression imputation.

Fits a regression (OLS/Logit) on one sample, then imputes a (possibly different) sample by binning the predicted yhat into percentile groups and drawing from the EMPIRICAL distribution of y within each bin - never a value predicted directly by the model. Unlike pmm/leaf error draws, nothing is donated from one recipient row to another; every draw comes from the model sample's own observed y values (or residuals, if draw_error=True), aggregated per bin.

The fitted model (coefficients, bin cutoffs, per-bin distribution) can be persisted to disk (path_save) and reloaded later (path_load/load_from_save) to impute a genuinely separate sample that was never in memory at the same time as the model sample - the "two-sample" in the name. Persistence is plain CSV/JSON files in a variable-specific folder (never pickle - not reliably portable across machines/environments/library versions - and never a bundled archive either, so every file can be opened directly without an extraction step), so model= must resolve to numeric-only predictors (a plain column list, or a formula with no C(...)/factor terms) - a fitted categorical encoding has no portable plain-text representation. Pre-encode any categorical predictor into numeric/dummy columns yourself first.

Parameters:

Name Type Description Default
model RegressionModel

OLS or Logit, by default RegressionModel.OLS (Logit if is_boolean=True and model isn't passed explicitly).

None
is_boolean bool

Whether impute_var is binary, by default False. Controls both the default model choice and the draw mechanism: bin-level P(y=1) + a Bernoulli draw, instead of a bin-level empirical quantile distribution.

False
bins int

Number of percentile bins to split predicted yhat into, by default 10 - each bin gets its own P(y=1) (boolean) or empirical quantile distribution (continuous) to draw from.

10
bin_by list[str] | str | None

Additional grouping columns - compute a separate distribution per (bin_by, prediction bin) cell instead of per prediction bin alone, by default None.

None
percentile_cuts list[float] | None

Explicit bin cut points (0-100 scale) instead of bins evenly spaced ones, by default None.

None
save_percentile_cuts bool

Freeze the bin cutoffs the first time they're computed under path_save, reusing the same cutoffs on every later call (e.g. across SRMI's repeated iterations) instead of recomputing them fresh each time - keeps disclosure-reviewed bin boundaries stable while the regression and distribution still refit every call. Independent of load_from_save, which freezes everything. By default False.

False
round_impute_var_digits int

Digits to DRB-round (see utilities.rounding.drb_round_table) the predicted score and impute_var to before computing bin cutoffs/distributions and before persisting them, by default 4.

4
continuous_qtiles_y_cuts list[float] | None

Quantile levels (0-1 scale) of the empirical y (or residual) distribution to compute per bin, later interpolated between to draw a value - see utilities.draw_from_quantiles. By default [0.1, 0.25, 0.5, 0.75, 0.9] if None.

None
continuous_qtiles_interpolate_by_bin bool

Compute the Census-style interpolation interval (Statistics(quantile_interpolated=True)) separately within each bin instead of once globally, by default False - more accurate when bins vary a lot in spread, at the cost of time.

False
min_n_x_var int

Minimum non-zero observations required per predictor to keep it in the model, by default 0 (no restriction) - for disclosure, same as _build_model_matrix's min_n_x_var.

0
draw_error bool

Draw from the bin's empirical distribution of residuals (yhat - y) and add the draw to yhat, instead of drawing directly from the bin's empirical distribution of y itself. Only meaningful when is_boolean=False. By default False.

False
random_share float

Fraction of the model sample to fit on, by default 1.0.

1.0
save_disclosure_support bool

Also persist a disclosure-review audit trail alongside the model (only when path_save is set): the exact record count in every (bin_by, prediction bin) cell the distribution was built from, and how many model+recipient records land exactly on each bin cutoff (relevant when the underlying data has heaping/rounding). Never affects the fitted model, the cutoffs, or the draw itself - purely an extra file for a human reviewer to check cell sizes with. By default False - these are real record counts, so leave this off unless someone is actually going to review them.

False
path_save str

Directory to persist the fitted model/cutoffs/distribution to (as plain files under f"{path_save}/{impute_var}/"), by default "" (don't persist).

''
path_load str

Directory to load a previously-persisted model from when load_from_save=True, by default "" (falls back to path_save).

''
load_from_save bool

Skip fitting entirely and impute purely from a previously persisted model (path_load/path_save) - the model sample doesn't need to be present in this run at all. By default False.

False

Returns:

Type Description
dict

Two-sample regression parameter dictionary

XGBoost staticmethod

XGBoost(
    parameters: dict | None = None,
    error: ErrorDraw = ErrorDraw.pmm,
    random_share: float = 1.0,
    cv_folds: int = 0,
    categorical_feature: list[str] | str | None = None,
    group_levels: list[str] | None = None,
    group_shrinkage_k: float = 10.0,
    parameters_pmm: dict = None,
    tune: bool = False,
    tuner=None,
) -> dict

Parameters for XGBoost-based imputation (mean regression only - see Parameters._tabular_ml_params's cv_folds docstring for why an honest donor-pool prediction matters more for flexible models like this one).

Parameters:

Name Type Description Default
parameters dict

Keyword arguments passed straight through to xgboost.XGBRegressor(**parameters), by default None (XGBRegressor's own defaults).

None
error ErrorDraw

Method for drawing errors - see Regression()'s error docstring, by default ErrorDraw.pmm. ErrorDraw.leaf is also available here (unlike Regression()) - donates by leaf co-occurrence in the fitted tree ensemble instead of PMM's knearest-on-scalar-yhat, the same mechanism Multinomial() uses for its RandomForestClassifier donor pool. Requires the fitted model to expose .apply() or .calc_leaf_indexes() - see ErrorDraw.leaf's own docstring.

pmm
random_share float

Fraction of data to use for fitting, by default 1.0.

1.0
cv_folds int

See Parameters._tabular_ml_params's cv_folds docstring, by default 0 (off).

0
categorical_feature list[str] | str | None

Predictor column names to treat as native categoricals (XGBoost's own histogram-based categorical splits, not one-hot encoding). Works with either form of the Variable's model= - a plain column list, or an R-style formula string as long as these columns are either left out of the formula entirely or referenced there only as a plain numeric term (see Variable._validate_estimator_available) - a bare reference to a non-numeric-dtype column in the formula still gets auto one-hot-encoded there regardless of this setting, which is what that validation catches. By default None (no native categoricals).

None
group_levels list[str] | None

See Regression()'s group_levels docstring, by default None (off).

None
group_shrinkage_k float

See Regression()'s group_shrinkage_k docstring, by default 10.0.

10.0
parameters_pmm dict

PMM parameters if using PMM error drawing, by default None.

None
tune bool

Whether to tune hyperparameters for this run, by default False. No effect if tuner is None. See utilities.tuning.Tuner - SRMI runs a Tuner.run_estimator() pass (over tuner's own HyperparameterSpace, against the constructed XGBRegressor as the template estimator) before the SRMI run starts.

False
tuner Tuner

Hyperparameter tuner - also owns whether/where tuned parameters get cached to disk (Tuner's own path_save_dir/overwrite, since the same tuner is typically reused across several variables, even across different modeltypes), by default None.

None

Returns:

Type Description
dict

XGBoost parameter dictionary

pmm staticmethod

pmm(
    knearest: int = 10,
    model: RegressionModel = RegressionModel.OLS,
    donate_list: list[str] = None,
    winsor: tuple[float, float] = [0, 1],
    donate_by: list[str] | str | None = None,
) -> dict

Parameters for predictive mean matching imputation.

Parameters:

Name Type Description Default
knearest int

Number of nearest neighbors for matching, by default 10

10
model RegressionModel

Regression model type, by default RegressionModel.OLS

OLS
donate_list list[str]

Additional variables to impute together, by default None i.e., you can predict earnings amount to find a donor, but then also impute hours worked and weeks worked with it.

None
winsor tuple[float, float]

Winsorization percentiles, by default [0, 1]

[0, 1]
donate_by list[str] | str | None

Grouping variables for donation, by default None

None

Returns:

Type Description
dict

PMM parameter dictionary

Selection

Bases: Serializable

Handles variable selection for imputation models.

Supports various selection methods including LASSO, stepwise selection, and custom functions to reduce model dimensionality.

Parameters:

Name Type Description Default
method Method

Selection method to use, by default Method.No

No
parameters dict

Method-specific parameters, by default None

None
select_within_by bool

Whether to run selection within each by-group, by default True

True
function callable

Custom selection function, by default None

None
Source code in src/survey_kit/imputation/selection.py
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
class Selection(Serializable):
    """
    Handles variable selection for imputation models.

    Supports various selection methods including LASSO, stepwise selection,
    and custom functions to reduce model dimensionality.

    Parameters
    ----------
    method : Selection.Method, optional
        Selection method to use, by default Method.No
    parameters : dict, optional
        Method-specific parameters, by default None
    select_within_by : bool, optional
        Whether to run selection within each by-group, by default True
    function : callable, optional
        Custom selection function, by default None
    """

    #   List of acceptable methods
    #       Commented out if not yet implemented
    class Method(Enum):
        No = 0
        Stepwise = 1
        LASSO = 2
        #   LightGBM = 3
        Custom = 4

    class Parameters:
        def lasso(
            nfolds: int = 5,
            type_measure: str = "default",
            include_base_with_interaction: bool = True,
            winsorize: tuple[float, float] | None = None,
            continuous: bool = False,
            binomial: bool = False,
            missing_dummies: bool = True,
            optimal_lambda: float = None,
            optimal_lambda_from_pre: bool = True,
            scale_lambda: float = 1.0,
        ) -> dict:
            """
            Parameters for LASSO variable selection method.

            Parameters
            ----------
            nfolds : int, default=5
                Number of folds for cross-validation during LASSO regression. Used to
                determine optimal lambda value through k-fold cross-validation.
            type_measure : str, default="default"
                Type of measure to use for cross-validation error. Determines how model
                performance is evaluated during lambda selection. Options are "default",
                "mse", "deviance", "class", "auc", "mae".
            include_base_with_interaction : bool, default=True
                Whether to include base variables when interaction terms are selected.
                If True, base variables are automatically included when their interactions
                are selected by LASSO.
            winsorize : tuple[float, float] or None, default=None
                Percentile bounds for winsorizing the dependent variable. Tuple of
                (low_percentile, high_percentile) to cap extreme values. None means
                no winsorization.
            continuous : bool, default=False
                Force treatment of dependent variable as continuous, overriding
                automatic detection.
            binomial : bool, default=False
                Force treatment of dependent variable as binomial, overriding
                automatic detection.
            missing_dummies : bool, default=True
                Whether to create dummy variables for missing values in predictor variables.
            optimal_lambda : float or None, default=None
                Pre-specified optimal lambda value for LASSO regularization. If None,
                optimal lambda will be determined through cross-validation.
            optimal_lambda_from_pre : bool, default=True
                Whether to use optimal lambda from a previous preselection step.
            scale_lambda : float, default=1.0
                Scaling factor applied to the optimal lambda value. Values > 1 make
                regularization stronger (fewer variables selected), values < 1 make
                it weaker (more variables selected).

            Returns
            -------
            dict
                Dictionary containing all parameter values for LASSO selection.

            Raises
            ------
            Exception
                When type_measure is not one of the acceptable values.
            """

            arguments = deepcopy(locals())

            type_measure_acceptable = [
                "default",
                "mse",
                "deviance",
                "class",
                "auc",
                "mae",
            ]
            #   Error checking
            if type_measure not in type_measure_acceptable:
                message = f"type_measure options are {type_measure_acceptable} (passed {type_measure})"
                logger.error(message)
                raise Exception(message)

            return arguments

        def lightgbm():
            return {}

        def stepwise(
            nfolds: int = 5,
            scoring: str = "neg_mean_squared_error",
            # Include base variables with interactions
            include_base_with_interaction: bool = True,
            # winsorize dependent on selection [low ptile, high ptile]
            winsorize: tuple[float, float] | None = None,
            # Force use of one or the other, otherwise, defaults based on data
            missing_dummies: bool = True,
            min_features_to_select: int = 10,
        ):
            """
            Parameters for stepwise variable selection method using Recursive Feature
            Elimination with Cross-Validation (RFECV).

            Parameters
            ----------
            nfolds : int, default=5
                Number of folds for cross-validation during stepwise selection. Used
                in RFECV to evaluate feature importance.
            scoring : str, default="neg_mean_squared_error"
                Scoring metric used to evaluate model performance during cross-validation.
                Should be a valid sklearn scoring parameter.
            include_base_with_interaction : bool, default=True
                Whether to include base variables when interaction terms are selected.
                If True, base variables are automatically included when their interactions
                are selected.
            winsorize : tuple[float, float] or None, default=None
                Percentile bounds for winsorizing the dependent variable. Tuple of
                (low_percentile, high_percentile) to cap extreme values. None means
                no winsorization.
            missing_dummies : bool, default=True
                Whether to create dummy variables for missing values in predictor variables.
            min_features_to_select : int, default=10
                Minimum number of features that must be selected by the stepwise procedure.
                Prevents over-reduction of the feature set.

            Returns
            -------
            dict
                Dictionary containing all parameter values for stepwise selection.
            """
            return deepcopy(locals())

    def __init__(
        self,
        method: Selection.Method = Method.No,
        parameters: dict = None,
        select_within_by: bool = True,
        function=None,
    ):
        """
        This class handles any variable selection steps needed to reduce
        the dimensionality of the imputation model

        Parameters
        ----------
        method : Selection.Method, optional
            Enumeration for existing selectio method. The default is Method.No (no selection).
        parameters : dict, optional
            Method-specific dictionary of parameters. The default is None.
        select_within_by : bool, optional
            During imputation, if True, run for each by group,
            if False, run once before all of them
        function: function handle, optional
            Custom selection function - if you want to use your own variable selection approach
                Function arguments need to be:
                df:IntoFrameT,
                y:str,
                formula:str,
                weight:str,
                sub_log:logging|None,
                parameters:dict,
                preselection:bool
                and it should return the selected model as a formula string
                (or "" to keep the original formula unchanged)

        Returns
        -------
        None.

        """

        if parameters is None:
            if method == Selection.Method.LASSO:
                parameters = Selection.Parameters.lasso()
            elif method == Selection.Method.Stepwise:
                parameters = Selection.Parameters.stepwise()
            #   elif method == Selection.Method.LightGBM:
            #       parameters = Selection.Parameters.lightgbm()
            else:
                parameters = {}

        if function is not None:
            method = Selection.Method.Custom
        self.method = method
        self.parameters = parameters
        self.function = function
        self.select_within_by = select_within_by

        #   Whether this Selection instance is being used as a Variable's
        #   preselection (Variable.__init__ flips this to True when it is) -
        #   defaulted here so Custom selection's run() can always read it,
        #   even when this instance is used as the main selection instead.
        self.preselection = False

    def run(
        self,
        df: IntoFrameT,
        y: str,
        formula: str,
        weight: str = "",
        sub_log: logging | None = None,
    ):
        arguments = {"df": df, "y": y, "formula": formula, "weight": weight, "sub_log": sub_log}

        delegate = None
        if self.method == Selection.Method.LASSO:
            delegate = self.lasso
        elif self.method == Selection.Method.Stepwise:
            delegate = self.stepwise
        # elif self.method == Selection.Method.LightGBM:
        #     delegate = self.lightgbm
        elif self.method == Selection.Method.Custom:
            delegate = self.function
            arguments["parameters"] = self.parameters
            arguments["preselection"] = self.preselection

        #   Call the actual selection function
        if delegate is None:
            return ""
        else:
            return delegate(**arguments)

    #   Selection methods - must have this signature/arguments
    def lasso(
        self,
        df: IntoFrameT,
        y: str,
        formula: str,
        weight: str = "",
        sub_log: logging | None = None,
    ):
        #   Save the original formula to compare to later
        formula_in = formula

        if sub_log is None:
            sub_log = logger

        #   Assign the parameter values
        nfolds = self.parameters["nfolds"]
        type_measure = self.parameters["type_measure"]
        winsorize = self.parameters["winsorize"]
        continuous = self.parameters["continuous"]
        binomial = self.parameters["binomial"]
        missing_dummies = self.parameters["missing_dummies"]
        optimal_lambda = self.parameters["optimal_lambda"]
        scale_lambda = self.parameters["scale_lambda"]

        #   collinearity_tolerance = self.parameters["collinearity_tolerance"]
        #   collinearity_tolerance_post = self.parameters["collinearity_tolerance_post"]

        include_base_with_interaction = self.parameters["include_base_with_interaction"]

        df = nw.from_native(df).filter(~nw.col(y).is_null()).to_native()

        if missing_dummies:
            [df, formula, missing_dummies] = Selection._add_missing_dummy(
                df=df, y=y, formula=formula
            )

        #   TODO - reimplement
        if winsorize is not None:
            df = winsorize_by_percentiles(df=df, percentiles=winsorize, columns=y)

        # matrix_y = as_matrix(PolarsToR(df_y))
        # del df_y

        #   TODO - Pre screen collinear variables?
        #   if collinearity_tolerance > 0:
        #       pass

        # r_args = {"y":matrix_y,
        #           "x":matrix_x,
        #           "parallel":True,
        #           "type.measure":type_measure,
        #           "alpha":1}

        # if weight != "":
        #     r_args["weights"] = as_matrix(PolarsToR(df.select(weight)))
        # if binomial:
        #     r_args["family"] = "binomial"
        # else:
        #     r_args["family"] = "gaussian"
        # if optimal_lambda is None:
        #     #   Get the lambda value
        #     r_args_crossval = deepcopy(r_args)
        #     r_args_crossval["nfolds"] = nfolds

        #     crossval = do_call(glmnet.cv_glmnet,dict_to_r_list(r_args_crossval))
        #     RInterop.set_item("crossval", crossval)
        #     optimal_lambda = RToPythonVariable(RInterop.get_item("crossval$lambda.min"))

        #     #   Assign it to the Selection object
        #     self.parameters["optimal_lambda"] = optimal_lambda

        #     #   Clean up
        #     RInterop.ReleaseMemory(remove_vars=["crossval"])

        #   With the optimal lambda set, run the lasso

        lm = rep_lasso(
            df=df,
            y=y,
            formula=formula,
            weight=weight,
            nfolds=nfolds,
            optimal_lambda=optimal_lambda,
        )

        if optimal_lambda is None:
            lm.find_optimal_lambda()
            optimal_lambda = lm.optimal_lambda

        if scale_lambda != 1:
            sub_log.info(f"         Scaling lambda by {scale_lambda}")
        lm.optimal_lambda = optimal_lambda * scale_lambda
        #   vars_kept = []
        vars_kept = lm.run()
        # try:

        #     # RInterop.set_item("fit", fit)
        #     # coefficients = RInterop.get_item("coef(fit)")
        #     # coefficients = RToPolars(as_data_frame(as_matrix(coefficients)),
        #     #                          KeepRowNames=True)
        #     # #   df_out = RToPolars(as_data_frame(tidy(fit)))
        #     # coefficients = coefficients.rename({
        #     #         coefficients.columns[0]:"Variable",
        #     #         coefficients.columns[1]:"Coefficient"
        #     #     })

        #     # #   Drop the empty coefficients
        #     # coefficients = coefficients.filter(pl.col("Coefficient") != 0)\
        #     #                            .filter(pl.col("Variable") != "(Intercept)")

        #     # vars_kept = coefficients["Variable"].to_list()
        # except Exception as error:
        #     sub_log.info(f"Error in glmnet call: {error}")

        if len(vars_kept) == 0:
            #   Are there any interactions, try without them
            [formula_nointeractions, any_dropped] = FormulaBuilder.exclude_interactions(
                formula=formula_in
            )

            if any_dropped:
                #   Changed formula, try lasso again
                output = self.lasso(
                    df=df,
                    y=y,
                    formula=formula_nointeractions,
                    weight=weight,
                    sub_log=sub_log,
                )
            else:
                #   Nothing changed, try stepwise
                output = self.stepwise(
                    df=df,
                    y=y,
                    formula=formula_nointeractions,
                    weight=weight,
                    sub_log=sub_log,
                )

            #   Return whatever we get - this runs recursively
            return output

        #   Clear the missingness dummies from the results
        if missing_dummies:
            vars_kept = [
                vari
                for vari in vars_kept
                if not vari.startswith("___missing___dummy___")
            ]

        #   Convert the final results into a formula
        fb = FormulaBuilder(df=df)
        fb.formula = formula_in
        fb.match_formula_to_columns(columns=vars_kept)

        if include_base_with_interaction:
            #   Get any base variables that go with the interaction
            fb.add_base_from_interactions()

        sub_log.info(f"         Selected model: ~{fb.rhs()}")
        return f"~{fb.rhs()}"

    def lightgbm(self, df: IntoFrameT, y: str, formula: str, weight: str = ""):
        pass

    def stepwise(
        self,
        df: IntoFrameT,
        y: str,
        formula: str,
        weight: str = "",
        sub_log: logging | None = None,
    ):
        #   Save the original formula to compare to later
        formula_in = formula

        if sub_log is None:
            sub_log = logger

        nfolds = self.parameters["nfolds"]

        if "scoring" in self.parameters.keys():
            scoring = self.parameters["scoring"]
        else:
            scoring = "neg_mean_squared_error"

        if "min_features_to_select" in self.parameters.keys():
            min_features_to_select = self.parameters["min_features_to_select"]
        else:
            min_features_to_select = 1

        winsorize = self.parameters["winsorize"]
        missing_dummies = self.parameters["missing_dummies"]
        include_base_with_interaction = self.parameters["include_base_with_interaction"]

        if weight != "":
            #   RFECV.fit() doesn't accept sample_weight without opting into
            #   sklearn's metadata-routing config (set_config(enable_metadata_routing=True)
            #   plus set_fit_request/set_score_request on the estimator and scorer) -
            #   not worth the added complexity/fragility for a selection method that
            #   isn't the primary one in use. Logging so this doesn't fail silently.
            sub_log.warning(
                f"Stepwise selection does not support sample weights - "
                f"running unweighted despite weight='{weight}' being set."
            )

        df = nw.from_native(df).filter(~nw.col(y).is_null()).to_native()
        if missing_dummies:
            [df, formula, missing_dummies] = Selection._add_missing_dummy(
                df=df, y=y, formula=formula
            )

        #   Imported here (not at module level) since sklearn's base import is
        #   ~450ms and Stepwise is the only Selection method that needs it -
        #   LASSO/HotDeck/StatMatch/LightGBM runs never call this and
        #   shouldn't pay for it.
        from sklearn.feature_selection import RFECV
        from sklearn.linear_model import LinearRegression

        estimator = LinearRegression()
        selector = RFECV(
            estimator,
            step=1,
            cv=nfolds,
            scoring=scoring,
            min_features_to_select=min_features_to_select,
        )

        df_x = get_model_frame(formula, df)

        df_y = nw.from_native(df).select(y).to_native()
        if winsorize is not None:
            df_y = winsorize_by_percentiles(df=df_y, percentiles=winsorize, columns=y)
        selector = selector.fit(
            nw.from_native(df_x).to_numpy(), nw.from_native(df_y).to_numpy()
        )

        vars_kept = []
        columns = nw.from_native(df_x).lazy().collect_schema().names()
        for i in range(0, len(columns)):
            if selector.support_[i]:
                vars_kept.append(columns[i])

        #   Clear the missingness dummies from the results
        if missing_dummies:
            vars_kept = [
                vari
                for vari in vars_kept
                if not vari.startswith("___missing___dummy___")
            ]

        #   Convert the final results into a formula
        fb = FormulaBuilder(df=df)
        fb.formula = formula_in
        fb.match_formula_to_columns(columns=vars_kept)

        if include_base_with_interaction:
            #   Get any base variables that go with the interaction
            fb.add_base_from_interactions()

        sub_log.info(f"         Selected model: ~{fb.rhs()}")
        return f"~{fb.rhs()}"

    def _add_missing_dummy(
        df: IntoFrameT, y: str, formula: str
    ) -> tuple[IntoFrameT, str, bool]:
        missing_dummies = []
        missing_recodes = []

        fb = FormulaBuilder(df=df)
        fb.formula = formula
        cols_to_check = [coli for coli in fb.columns if coli != y]

        #   One pass to find which columns have any missing values, instead
        #       of a separate filter/materialize per column.
        any_missing = {}
        if len(cols_to_check) > 0:
            any_missing = list(
                nw.from_native(df)
                .select([nw.col(coli).is_null().any().alias(coli) for coli in cols_to_check])
                .lazy()
                .collect()
                .iter_rows(named=True)
            )[0]

        for coli in cols_to_check:
            if any_missing.get(coli, False):
                missing_dummies.append(
                    nw.when(nw.col(coli).is_null())
                    .then(nw.lit(True))
                    .otherwise(nw.lit(False))
                    .alias(f"___missing___dummy___{coli}")
                )
                missing_recodes.append(
                    nw.when(nw.col(coli).is_null())
                    .then(nw.lit(0))
                    .otherwise(nw.col(coli))
                    .alias(coli)
                )

        if len(missing_dummies) > 0:
            df = (
                nw.from_native(df)
                .with_columns(missing_dummies + missing_recodes)
                .to_native()
            )

            #   Add these variables to the formula
            #       Update the formula to the new dataframe so it can find the variables
            fb.df = df
            fb.continuous(columns="___missing___dummy___*")

            return (df, formula, True)
        else:
            #   None missing
            return (df, formula, False)

Parameters

Source code in src/survey_kit/imputation/selection.py
class Parameters:
    def lasso(
        nfolds: int = 5,
        type_measure: str = "default",
        include_base_with_interaction: bool = True,
        winsorize: tuple[float, float] | None = None,
        continuous: bool = False,
        binomial: bool = False,
        missing_dummies: bool = True,
        optimal_lambda: float = None,
        optimal_lambda_from_pre: bool = True,
        scale_lambda: float = 1.0,
    ) -> dict:
        """
        Parameters for LASSO variable selection method.

        Parameters
        ----------
        nfolds : int, default=5
            Number of folds for cross-validation during LASSO regression. Used to
            determine optimal lambda value through k-fold cross-validation.
        type_measure : str, default="default"
            Type of measure to use for cross-validation error. Determines how model
            performance is evaluated during lambda selection. Options are "default",
            "mse", "deviance", "class", "auc", "mae".
        include_base_with_interaction : bool, default=True
            Whether to include base variables when interaction terms are selected.
            If True, base variables are automatically included when their interactions
            are selected by LASSO.
        winsorize : tuple[float, float] or None, default=None
            Percentile bounds for winsorizing the dependent variable. Tuple of
            (low_percentile, high_percentile) to cap extreme values. None means
            no winsorization.
        continuous : bool, default=False
            Force treatment of dependent variable as continuous, overriding
            automatic detection.
        binomial : bool, default=False
            Force treatment of dependent variable as binomial, overriding
            automatic detection.
        missing_dummies : bool, default=True
            Whether to create dummy variables for missing values in predictor variables.
        optimal_lambda : float or None, default=None
            Pre-specified optimal lambda value for LASSO regularization. If None,
            optimal lambda will be determined through cross-validation.
        optimal_lambda_from_pre : bool, default=True
            Whether to use optimal lambda from a previous preselection step.
        scale_lambda : float, default=1.0
            Scaling factor applied to the optimal lambda value. Values > 1 make
            regularization stronger (fewer variables selected), values < 1 make
            it weaker (more variables selected).

        Returns
        -------
        dict
            Dictionary containing all parameter values for LASSO selection.

        Raises
        ------
        Exception
            When type_measure is not one of the acceptable values.
        """

        arguments = deepcopy(locals())

        type_measure_acceptable = [
            "default",
            "mse",
            "deviance",
            "class",
            "auc",
            "mae",
        ]
        #   Error checking
        if type_measure not in type_measure_acceptable:
            message = f"type_measure options are {type_measure_acceptable} (passed {type_measure})"
            logger.error(message)
            raise Exception(message)

        return arguments

    def lightgbm():
        return {}

    def stepwise(
        nfolds: int = 5,
        scoring: str = "neg_mean_squared_error",
        # Include base variables with interactions
        include_base_with_interaction: bool = True,
        # winsorize dependent on selection [low ptile, high ptile]
        winsorize: tuple[float, float] | None = None,
        # Force use of one or the other, otherwise, defaults based on data
        missing_dummies: bool = True,
        min_features_to_select: int = 10,
    ):
        """
        Parameters for stepwise variable selection method using Recursive Feature
        Elimination with Cross-Validation (RFECV).

        Parameters
        ----------
        nfolds : int, default=5
            Number of folds for cross-validation during stepwise selection. Used
            in RFECV to evaluate feature importance.
        scoring : str, default="neg_mean_squared_error"
            Scoring metric used to evaluate model performance during cross-validation.
            Should be a valid sklearn scoring parameter.
        include_base_with_interaction : bool, default=True
            Whether to include base variables when interaction terms are selected.
            If True, base variables are automatically included when their interactions
            are selected.
        winsorize : tuple[float, float] or None, default=None
            Percentile bounds for winsorizing the dependent variable. Tuple of
            (low_percentile, high_percentile) to cap extreme values. None means
            no winsorization.
        missing_dummies : bool, default=True
            Whether to create dummy variables for missing values in predictor variables.
        min_features_to_select : int, default=10
            Minimum number of features that must be selected by the stepwise procedure.
            Prevents over-reduction of the feature set.

        Returns
        -------
        dict
            Dictionary containing all parameter values for stepwise selection.
        """
        return deepcopy(locals())

lasso

lasso(
    nfolds: int = 5,
    type_measure: str = "default",
    include_base_with_interaction: bool = True,
    winsorize: tuple[float, float] | None = None,
    continuous: bool = False,
    binomial: bool = False,
    missing_dummies: bool = True,
    optimal_lambda: float = None,
    optimal_lambda_from_pre: bool = True,
    scale_lambda: float = 1.0,
) -> dict

Parameters for LASSO variable selection method.

Parameters:

Name Type Description Default
nfolds int

Number of folds for cross-validation during LASSO regression. Used to determine optimal lambda value through k-fold cross-validation.

5
type_measure str

Type of measure to use for cross-validation error. Determines how model performance is evaluated during lambda selection. Options are "default", "mse", "deviance", "class", "auc", "mae".

"default"
include_base_with_interaction bool

Whether to include base variables when interaction terms are selected. If True, base variables are automatically included when their interactions are selected by LASSO.

True
winsorize tuple[float, float] or None

Percentile bounds for winsorizing the dependent variable. Tuple of (low_percentile, high_percentile) to cap extreme values. None means no winsorization.

None
continuous bool

Force treatment of dependent variable as continuous, overriding automatic detection.

False
binomial bool

Force treatment of dependent variable as binomial, overriding automatic detection.

False
missing_dummies bool

Whether to create dummy variables for missing values in predictor variables.

True
optimal_lambda float or None

Pre-specified optimal lambda value for LASSO regularization. If None, optimal lambda will be determined through cross-validation.

None
optimal_lambda_from_pre bool

Whether to use optimal lambda from a previous preselection step.

True
scale_lambda float

Scaling factor applied to the optimal lambda value. Values > 1 make regularization stronger (fewer variables selected), values < 1 make it weaker (more variables selected).

1.0

Returns:

Type Description
dict

Dictionary containing all parameter values for LASSO selection.

Raises:

Type Description
Exception

When type_measure is not one of the acceptable values.

Source code in src/survey_kit/imputation/selection.py
def lasso(
    nfolds: int = 5,
    type_measure: str = "default",
    include_base_with_interaction: bool = True,
    winsorize: tuple[float, float] | None = None,
    continuous: bool = False,
    binomial: bool = False,
    missing_dummies: bool = True,
    optimal_lambda: float = None,
    optimal_lambda_from_pre: bool = True,
    scale_lambda: float = 1.0,
) -> dict:
    """
    Parameters for LASSO variable selection method.

    Parameters
    ----------
    nfolds : int, default=5
        Number of folds for cross-validation during LASSO regression. Used to
        determine optimal lambda value through k-fold cross-validation.
    type_measure : str, default="default"
        Type of measure to use for cross-validation error. Determines how model
        performance is evaluated during lambda selection. Options are "default",
        "mse", "deviance", "class", "auc", "mae".
    include_base_with_interaction : bool, default=True
        Whether to include base variables when interaction terms are selected.
        If True, base variables are automatically included when their interactions
        are selected by LASSO.
    winsorize : tuple[float, float] or None, default=None
        Percentile bounds for winsorizing the dependent variable. Tuple of
        (low_percentile, high_percentile) to cap extreme values. None means
        no winsorization.
    continuous : bool, default=False
        Force treatment of dependent variable as continuous, overriding
        automatic detection.
    binomial : bool, default=False
        Force treatment of dependent variable as binomial, overriding
        automatic detection.
    missing_dummies : bool, default=True
        Whether to create dummy variables for missing values in predictor variables.
    optimal_lambda : float or None, default=None
        Pre-specified optimal lambda value for LASSO regularization. If None,
        optimal lambda will be determined through cross-validation.
    optimal_lambda_from_pre : bool, default=True
        Whether to use optimal lambda from a previous preselection step.
    scale_lambda : float, default=1.0
        Scaling factor applied to the optimal lambda value. Values > 1 make
        regularization stronger (fewer variables selected), values < 1 make
        it weaker (more variables selected).

    Returns
    -------
    dict
        Dictionary containing all parameter values for LASSO selection.

    Raises
    ------
    Exception
        When type_measure is not one of the acceptable values.
    """

    arguments = deepcopy(locals())

    type_measure_acceptable = [
        "default",
        "mse",
        "deviance",
        "class",
        "auc",
        "mae",
    ]
    #   Error checking
    if type_measure not in type_measure_acceptable:
        message = f"type_measure options are {type_measure_acceptable} (passed {type_measure})"
        logger.error(message)
        raise Exception(message)

    return arguments

stepwise

stepwise(
    nfolds: int = 5,
    scoring: str = "neg_mean_squared_error",
    include_base_with_interaction: bool = True,
    winsorize: tuple[float, float] | None = None,
    missing_dummies: bool = True,
    min_features_to_select: int = 10,
)

Parameters for stepwise variable selection method using Recursive Feature Elimination with Cross-Validation (RFECV).

Parameters:

Name Type Description Default
nfolds int

Number of folds for cross-validation during stepwise selection. Used in RFECV to evaluate feature importance.

5
scoring str

Scoring metric used to evaluate model performance during cross-validation. Should be a valid sklearn scoring parameter.

"neg_mean_squared_error"
include_base_with_interaction bool

Whether to include base variables when interaction terms are selected. If True, base variables are automatically included when their interactions are selected.

True
winsorize tuple[float, float] or None

Percentile bounds for winsorizing the dependent variable. Tuple of (low_percentile, high_percentile) to cap extreme values. None means no winsorization.

None
missing_dummies bool

Whether to create dummy variables for missing values in predictor variables.

True
min_features_to_select int

Minimum number of features that must be selected by the stepwise procedure. Prevents over-reduction of the feature set.

10

Returns:

Type Description
dict

Dictionary containing all parameter values for stepwise selection.

Source code in src/survey_kit/imputation/selection.py
def stepwise(
    nfolds: int = 5,
    scoring: str = "neg_mean_squared_error",
    # Include base variables with interactions
    include_base_with_interaction: bool = True,
    # winsorize dependent on selection [low ptile, high ptile]
    winsorize: tuple[float, float] | None = None,
    # Force use of one or the other, otherwise, defaults based on data
    missing_dummies: bool = True,
    min_features_to_select: int = 10,
):
    """
    Parameters for stepwise variable selection method using Recursive Feature
    Elimination with Cross-Validation (RFECV).

    Parameters
    ----------
    nfolds : int, default=5
        Number of folds for cross-validation during stepwise selection. Used
        in RFECV to evaluate feature importance.
    scoring : str, default="neg_mean_squared_error"
        Scoring metric used to evaluate model performance during cross-validation.
        Should be a valid sklearn scoring parameter.
    include_base_with_interaction : bool, default=True
        Whether to include base variables when interaction terms are selected.
        If True, base variables are automatically included when their interactions
        are selected.
    winsorize : tuple[float, float] or None, default=None
        Percentile bounds for winsorizing the dependent variable. Tuple of
        (low_percentile, high_percentile) to cap extreme values. None means
        no winsorization.
    missing_dummies : bool, default=True
        Whether to create dummy variables for missing values in predictor variables.
    min_features_to_select : int, default=10
        Minimum number of features that must be selected by the stepwise procedure.
        Prevents over-reduction of the feature set.

    Returns
    -------
    dict
        Dictionary containing all parameter values for stepwise selection.
    """
    return deepcopy(locals())

tuning

Categorical dataclass

A hyperparameter chosen from a fixed set of values.

FloatRange dataclass

A float hyperparameter, uniform (or log-uniform) between low/high.

HyperparameterSpace

A named, typed hyperparameter search space - replaces a bare {"num_leaves": [2, 256]} dict with self-documenting IntRange/FloatRange/ Categorical fields, e.g.:

HyperparameterSpace(
    num_leaves=IntRange(2, 256),
    learning_rate=FloatRange(1e-3, 0.3, log=True),
    boosting=Categorical(["gbdt", "dart"]),
)

Extensible: any object with a .suggest(trial, name) method can be used as a field value - a new distribution type is just a new small class, not a change to HyperparameterSpace or Tuner.

IntRange dataclass

An integer hyperparameter, uniform (or log-uniform) between low/high.

Objective

Bases: Enum

A scoring function for comparing a model's raw predictions against holdout labels - used by Tuner.run_lightgbm (LightGBM's own .predict() doesn't go through an sklearn Estimator, so it can't use run_estimator's sklearn scoring= string instead). Values must NOT be plain functions - Enum silently treats function-valued (or otherwise descriptor-like, e.g. functools.partial on newer Python) class attributes as methods rather than registering them as real members, so string values are used here instead and the actual scoring function is looked up via _objective_scorers() below.

Tuner

One tuner class for every imputation model family - the same instance is meant to be reused across many Variables (e.g. via SRMI.Defaults), each getting its own independent search and its own result:

tuner = Tuner(
    space=HyperparameterSpace(num_leaves=IntRange(2, 256), ...),
    n_trials=50,
    path_save_dir="tuner_outputs",
    overwrite=True,
)
Parameters.LightGBM(tune=True, tuner=tuner, ...)      # var A
Parameters.LightGBM(tune=True, tuner=tuner, ...)      # var B, same tuner

Every routing call below (run/run_lightgbm/run_estimator) builds a FRESH optuna study - reusing one study across variables would silently mix each variable's (unrelated) trials into the same search, and study.best_params/best_value would then reflect whichever variable happened to score best overall, not the one you just tuned. Only space/n_trials/direction/seed/path_save_dir/overwrite are shared state; the study itself never is.

Three ways to use it, in increasing order of how much you write: - run_lightgbm(train_data, test_data, base_params) - built-in LightGBM routing (CV over lgb.basic.Dataset, scored via objective). - run_estimator(estimator, X, y, cv=..., scoring=..., fit_params=...) - built-in sklearn routing, argument names matching sklearn.model_selection's own (GridSearchCV-style: clone(estimator) .set_params(**trial_params) per trial). - run(score_fn) - the escape hatch for anything else: score_fn(params: dict) -> float is called once per trial with that trial's suggested hyperparameters; fit/score however you need to, inside your own closure. run_lightgbm/run_estimator are themselves just score_fn closures built for you and passed to this same method.

run

run(
    score_fn: Callable[[dict], float],
    direction: str | None = None,
) -> dict

The generic primitive every routing method (including your own) is built on: score_fn(params: dict) -> float is called once per trial. Runs a FRESH study every call, saves the result to path_save (if set) and returns study.best_params.

run_estimator

run_estimator(
    estimator,
    X,
    y,
    cv: int = 3,
    scoring: str = "neg_root_mean_squared_error",
    fit_params: dict | None = None,
    direction: str = "maximize",
) -> dict

sklearn-style CV, argument names matching sklearn.model_selection's own (GridSearchCV/cross_val_score: cv, scoring, fit_params) so this reads like ordinary sklearn code. For each trial, clones estimator (an unfitted, already-constructed sklearn-compatible estimator - the same clone()+set_params() idiom GridSearchCV itself uses), sets that trial's suggested hyperparameters on the clone, fits/scores it across cv folds (a fixed fold assignment reused across every trial, so score differences reflect hyperparameter differences, not fold-split noise), and returns the mean fold score.

Splitting is done natively in polars: X (and y and any per-row fit_params, e.g. sample_weight) are attached as columns on one combined frame alongside a random fold-id column, then partition_by(fold_id, include_key=False) both groups AND drops the fold-id column in one step - so it can never leak into X as a feature. This is this codebase's own native dataframe library (used everywhere else in survey_kit) rather than reaching for sklearn's private _safe_indexing (no backward-compatibility guarantee per its own docstring) or supporting pandas (not used anywhere in this codebase).

Parameters:

Name Type Description Default
estimator sklearn-compatible estimator

An unfitted, already-constructed instance (e.g. RandomForestRegressor()) - cloned fresh for every trial/fold, never fit in place.

required
X DataFrame | ndarray

Already numeric/model-matrix-ready training data - a polars DataFrame (as Impute._prepare_tuning_data produces - needed to keep XGBoost's/CatBoost's native categorical dtype intact through to .fit()) or a plain numpy array (wrapped into a polars DataFrame internally). No pandas.

required
y array - like

Target values, length matching X.

required
cv int

Number of cross-validation folds, by default 3.

3
scoring str

Any sklearn scoring string (sklearn.metrics.get_scorer_names()), by default "neg_root_mean_squared_error".

'neg_root_mean_squared_error'
fit_params dict

Extra keyword arguments passed to every fold's .fit() call (e.g. {"sample_weight": w}) - split to each fold the same way X/y are, if array-like of matching length; passed as-is otherwise. By default None.

None
direction str

"maximize" or "minimize" - by default "maximize" (matching sklearn's own scorer convention, where higher is better - hence "neg_..." for error metrics).

'maximize'

run_lightgbm

run_lightgbm(
    train_data,
    test_data,
    base_params: dict,
    objective: Objective | None = None,
) -> dict

LightGBM's own CV: for each trial, train on train_data (an already- built lgb.basic.Dataset) with base_params merged with that trial's suggested hyperparameters, and score the fitted model's predictions on test_data against test_data.label via objective (falls back to this tuner's own self.objective, set at construction, if not given here).

load_tuned_params

load_tuned_params(path: str) -> dict | None

Read back a hyperparameter dict previously saved by Tuner.run() (via path_save) - returns None if path is empty or nothing has been saved there yet (e.g. the very first run, before any tuning pass has completed). Used both by SRMI's own tune-before-run preprocessing (to check whether a fresh tuning pass is even needed) and by the actual per-iteration fit (to re-load the tuned values fresh from disk every time, rather than trying to carry a fitted-in-memory value across a deepcopy(variable) or a parallel worker process boundary).