Skip to content

Regression Adapters

Every adapter below returns the same normalized type - AdapterStats - so mi_ses_from_function treats any of them identically (it requires a StatCalculator, which AdapterStats is). Each also has a matching mi_ses_from_<package>(...) shortcut that runs it across multiple-imputation implicates directly, taking the adapter's own arguments as keywords instead of an arguments={} dict - see Using the Regression Adapters for a walkthrough.

AdapterStats

The shared return type - a StatCalculator subclass built directly from df_estimates/df_ses (plus df_vcov/df_tidy when available), with no raw microdata or replicate weights involved. Use this directly if you're wiring up your own delegate for a package with no named adapter below (see Rolling Your Own) - build one from your own df_estimates/df_ses and it's already a valid mi_ses_from_function delegate.

AdapterStats

Bases: StatCalculator

A StatCalculator built directly from an already-computed set of estimates + standard errors - e.g. a regression adapter's own analytic output (see survey_kit.statistics.adapters, every one of which returns this) - rather than from raw microdata. No replicate-weight recomputation at all; use StatCalculator.from_function instead if you want SEs derived from the spread across replicate weights rather than df_ses's own values.

Matches StatCalculator's own API and function surface exactly (comparisons, printing, save/load, and - since MultipleImputation's own combination code reads whatever a delegate returns generically - mi_ses_from_function/MultipleImputation work with this the same way they do a plain StatCalculator) since it IS one, just constructed differently - StatCalculator.copy() (which every chainable method starts from) preserves the real subclass rather than rebuilding a plain StatCalculator, so this stays an AdapterStats through filter()/select()/sort()/rename()/with_columns()/scale_by()/etc.

filter()/rename() are additionally extended here to keep df_vcov in sync - the base class's versions only ever touch df_estimates/df_ses/df_replicates (df_vcov didn't exist before this class), so a plain filter()/rename() would otherwise just drop it (with a warning) rather than correctly narrow/rename it. with_columns()/drop()/pipe() operate on value columns or arbitrary user logic, where there's no safe, generic way to know whether/how df_vcov/df_tidy (tied to one specific coefficient column, or to the source package's own unrelated shape) should follow along - rather than silently dropping either one, these raise ValueError instead when df_vcov/df_tidy is set, so a caller finds out immediately rather than discovering it missing later. Clear the one(s) you don't need first (e.g. obj.replicate_stats.df_vcov = None) to use these methods anyway. select() doesn't touch df_vcov/df_tidy at all (for the same reason as those three - not because it was overlooked) but also doesn't raise, since narrowing which estimate columns are kept has no "selected columns" concept for either one to begin with.

concat_with() also raises for df_tidy (same reasoning as above) and for df_vcov on a horizontal concat (it adds a new value column, breaking df_vcov's single-value-column precondition) - but on a vertical concat with both sides carrying a df_vcov over disjoint terms, it stacks them block-diagonally instead of dropping either one, logging a warning that cross-object covariance is assumed zero/unknown (the same independence assumption .compare() already makes between two separate objects).

.compare() (inherited from StatCalculator) computes a difference/ ratio SE from df_ses under an independence assumption between the two objects being compared, but always clears the result's df_vcov (it described the original fit's own term-by-term covariance, which doesn't carry over to a cross-object difference/ratio without more information than either object has).

Source code in src/survey_kit/statistics/adapter_stats.py
class AdapterStats(StatCalculator):
    """
    A StatCalculator built directly from an already-computed set of
    estimates + standard errors - e.g. a regression adapter's own
    analytic output (see survey_kit.statistics.adapters, every one of
    which returns this) - rather than from raw microdata. No
    replicate-weight recomputation at all; use
    StatCalculator.from_function instead if you want SEs derived from
    the spread across replicate weights rather than df_ses's own
    values.

    Matches StatCalculator's own API and function surface exactly
    (comparisons, printing, save/load, and - since MultipleImputation's
    own combination code reads whatever a delegate returns generically -
    mi_ses_from_function/MultipleImputation work with this the same way
    they do a plain StatCalculator) since it IS one, just constructed
    differently - StatCalculator.copy() (which every chainable method
    starts from) preserves the real subclass rather than rebuilding a
    plain StatCalculator, so this stays an AdapterStats through
    filter()/select()/sort()/rename()/with_columns()/scale_by()/etc.

    filter()/rename() are additionally extended here to keep df_vcov in
    sync - the base class's versions only ever touch
    df_estimates/df_ses/df_replicates (df_vcov didn't exist before this
    class), so a plain filter()/rename() would otherwise just drop it
    (with a warning) rather than correctly narrow/rename it.
    with_columns()/drop()/pipe() operate on value columns or arbitrary
    user logic, where there's no safe, generic way to know whether/how
    df_vcov/df_tidy (tied to one specific coefficient column, or to the
    source package's own unrelated shape) should follow along - rather
    than silently dropping either one, these raise ValueError instead
    when df_vcov/df_tidy is set, so a caller finds out immediately
    rather than discovering it missing later. Clear the one(s) you
    don't need first (e.g. `obj.replicate_stats.df_vcov = None`) to use
    these methods anyway. select() doesn't touch df_vcov/df_tidy at all
    (for the same reason as those three - not because it was
    overlooked) but also doesn't raise, since narrowing *which*
    estimate columns are kept has no "selected columns" concept for
    either one to begin with.

    concat_with() also raises for df_tidy (same reasoning as above) and
    for df_vcov on a horizontal concat (it adds a new value column,
    breaking df_vcov's single-value-column precondition) - but on a
    vertical concat with both sides carrying a df_vcov over disjoint
    terms, it stacks them block-diagonally instead of dropping either
    one, logging a warning that cross-object covariance is assumed
    zero/unknown (the same independence assumption .compare() already
    makes between two separate objects).

    .compare() (inherited from StatCalculator) computes a difference/
    ratio SE from df_ses under an independence assumption between the
    two objects being compared, but always clears the result's df_vcov
    (it described the original fit's own term-by-term covariance, which
    doesn't carry over to a cross-object difference/ratio without more
    information than either object has).
    """

    def __init__(
        self,
        df_estimates: IntoFrameT,
        df_ses: IntoFrameT,
        variable_ids: list[str] | str = "Variable",
        by: dict[str, list[str]] | list | None = None,
        df_vcov: IntoFrameT | None = None,
        df_tidy: IntoFrameT | None = None,
        display: bool = True,
        display_all_vars: bool = True,
        display_max_vars: int = 20,
        round_output: bool | int = True,
    ):
        """
        Parameters
        ----------
        df_estimates : IntoFrameT
            One row per estimate - variable_ids (+ any `by` columns)
            plus one column per estimated quantity (e.g. a regression's
            coefficients).
        df_ses : IntoFrameT
            Same shape as df_estimates - standard errors for each value.
        variable_ids : list[str] | str, optional
            Column name(s) identifying each row. Default "Variable"
            (matches every adapter in survey_kit.statistics.adapters,
            which all use this as their own join_on_name default).
        by : dict[str,list[str]] | list | None, optional
            Grouping columns already present in df_estimates/df_ses
            (e.g. if the estimator was run separately per group and the
            results concatenated) - same shape as StatCalculator's own
            `by`. Default None.
        df_vcov : IntoFrameT | None, optional
            Variance-covariance matrix, long/pairwise: for each id in
            variable_ids, two copies of that column suffixed "_1" and
            "_2" (row term, column term), plus one value column matching
            df_estimates' own single non-join_on column - only
            meaningful when df_estimates has exactly one such column
            (e.g. a regression coefficient table). This is what lets
            .compare() compute a correct joint SE between two correlated
            rows of this same fit (e.g. two coefficients), instead of
            assuming independence. Default None.
        df_tidy : IntoFrameT | None, optional
            The underlying package's own native summary table, kept
            as-is for reference - never used in any computation here,
            and not kept in sync by filter()/select() (its own id
            column(s), if any, aren't guaranteed to match variable_ids).
            Default None.
        display, display_all_vars, display_max_vars, round_output
            Same as StatCalculator's own constructor.
        """
        if isinstance(variable_ids, str):
            variable_ids = [variable_ids]

        super().__init__(
            df=None,
            by=by,
            display=display,
            display_all_vars=display_all_vars,
            display_max_vars=display_max_vars,
            round_output=round_output,
            calculate=False,
        )
        self.variable_ids = variable_ids
        self.replicate_stats = ReplicateStats(
            df_estimates=df_estimates,
            df_ses=df_ses,
            df_vcov=df_vcov,
            df_tidy=df_tidy,
        )
        self.df_estimates = self.round_results(df=self.df_estimates)
        self.df_ses = self.round_results(df=self.df_ses)

        if display:
            self.print()

    def print(
        self, round_output=None, estimates_per_page: int = 0, sub_log=None
    ) -> None:
        #   The base class's print() decides whether to show SEs by
        #   checking df_replicates (raw per-replicate draws) - AdapterStats
        #   never has those (its SEs come directly from the adapter, not
        #   from replicate weights), so that check would wrongly hide
        #   df_ses, which IS populated here. Check df_ses directly
        #   instead; _print_replicates itself only ever reads
        #   df_estimates/df_ses; despite the name, it never touches
        #   df_replicates.
        if self.df_ses is not None:
            self._print_replicates(
                round_output=round_output,
                estimates_per_page=estimates_per_page,
                sub_log=sub_log,
            )
        else:
            self._print_estimates(
                round_output=round_output,
                estimates_per_page=estimates_per_page,
                sub_log=sub_log,
            )

    def filter(self, filter_expr: nw.Expr) -> AdapterStats:
        #   ReplicateStats.filter() (via _invalidate_extras) already
        #   drops df_vcov itself - it has no generic way to reshape
        #   term-pair rows - logging a warning as it does. Capture the
        #   original here, before that happens, so it can be correctly
        #   re-narrowed afterward instead of lost.
        original_vcov = self.replicate_stats.df_vcov

        self = super().filter(filter_expr)

        if original_vcov is not None:
            self.replicate_stats.df_vcov = _vcov_semi_join(
                original_vcov, self.variable_ids, self.df_estimates
            )
        return self

    def select(
        self, select_expr: nw.Expr | str | list[str] | list[nw.Expr]
    ) -> AdapterStats:
        #   Column selection narrows which *estimate* columns remain -
        #   df_vcov describes the covariance of a single coefficient
        #   column (see its own docstring above) and df_tidy is an
        #   unrelated native snapshot, so neither one has a "selected
        #   columns" concept of its own to keep in sync here; only rows
        #   (filter(), above) can make them inconsistent.
        return super().select(select_expr)

    def rename(self, d_rename: dict[str, str]) -> AdapterStats:
        #   Same _invalidate_extras drop-and-warn as filter() - capture
        #   df_vcov first and re-derive it afterward, applying the same
        #   rename to whichever of its own columns are affected: an id
        #   column (variable_ids) renames both its "_1"/"_2" copies, the
        #   value column (if renamed) renames directly.
        original_vcov = self.replicate_stats.df_vcov

        self = super().rename(d_rename)

        if original_vcov is not None:
            vcov_cols = set(
                nw.from_native(original_vcov).lazy().collect_schema().names()
            )
            vcov_rename = {}
            for old, new in d_rename.items():
                if f"{old}_1" in vcov_cols:
                    vcov_rename[f"{old}_1"] = f"{new}_1"
                if f"{old}_2" in vcov_cols:
                    vcov_rename[f"{old}_2"] = f"{new}_2"
                if old in vcov_cols:
                    vcov_rename[old] = new

            self.replicate_stats.df_vcov = (
                nw.from_native(original_vcov).rename(vcov_rename).to_native()
                if vcov_rename
                else original_vcov
            )
        return self

    def with_columns(self, with_expr: nw.Expr | list[nw.Expr]) -> AdapterStats:
        self._raise_if_vcov_or_tidy("with_columns")
        return super().with_columns(with_expr)

    def drop(
        self, drop_expr: nw.Expr | list[nw.Expr] | str | list[str]
    ) -> AdapterStats:
        self._raise_if_vcov_or_tidy("drop")
        return super().drop(drop_expr)

    def pipe(self, function, *args, **kwargs) -> AdapterStats:
        self._raise_if_vcov_or_tidy("pipe")
        return super().pipe(function, *args, **kwargs)

    def concat_with(
        self, sc_concat: StatCalculator, how: str = "horizontal"
    ) -> AdapterStats:
        #   df_tidy never has a generic story here regardless of `how` -
        #   its own id column(s), if any, aren't guaranteed to match
        #   variable_ids (see the class docstring), so there's no way to
        #   tell which of its rows would even correspond to which
        #   post-concat row.
        other_replicate_stats = getattr(sc_concat, "replicate_stats", None)
        self_vcov = self.replicate_stats.df_vcov
        other_vcov = getattr(other_replicate_stats, "df_vcov", None)
        if self.replicate_stats.df_tidy is not None or (
            getattr(other_replicate_stats, "df_tidy", None) is not None
        ):
            raise ValueError(
                "AdapterStats.concat_with(): df_tidy can't be reshaped "
                "generically and won't be silently dropped - clear it on "
                "whichever side has it first (e.g. "
                "self.replicate_stats.df_tidy = None) if you don't need it."
            )

        if how == "horizontal":
            #   A horizontal concat adds a *new value column* from
            #   sc_concat - df_vcov only ever describes one value
            #   column's own covariance (see the class docstring), so
            #   even under a block-independence assumption there's no
            #   single-value-column schema left to put the result in.
            if self_vcov is not None or other_vcov is not None:
                raise ValueError(
                    "AdapterStats.concat_with(how='horizontal'): can't "
                    "carry df_vcov through a horizontal concat (it adds a "
                    "new value column, and df_vcov only describes one "
                    "value column's own covariance) - clear it on "
                    "whichever side has it first (e.g. "
                    "self.replicate_stats.df_vcov = None) if you don't "
                    "need it."
                )
            return super().concat_with(sc_concat, how=how)

        #   how == "vertical": stacking rows is compatible with df_vcov
        #   as long as both sides have one (or neither) and their terms
        #   are disjoint - concatenate the two long tables block-
        #   diagonally, with cross-covariance between a self term and a
        #   sc_concat term treated as unknown/zero (the same
        #   independence assumption StatCalculator.compare() already
        #   makes between two separate objects).
        if self_vcov is None and other_vcov is None:
            return super().concat_with(sc_concat, how=how)

        if (self_vcov is None) != (other_vcov is None):
            raise ValueError(
                "AdapterStats.concat_with(how='vertical'): only one side "
                "has a df_vcov - stacking would either silently drop it "
                "or leave the other side's terms with no covariance "
                "information at all. Clear it on whichever side has it "
                "(e.g. self.replicate_stats.df_vcov = None) if you don't "
                "need it."
            )

        if self.variable_ids != sc_concat.variable_ids:
            raise ValueError(
                "AdapterStats.concat_with(how='vertical'): self and "
                "sc_concat have different variable_ids - can't line up "
                "df_vcov's {id}_1/{id}_2 columns between them."
            )

        self_terms = (
            nw.from_native(self.df_estimates).lazy().select(self.variable_ids).unique()
        )
        other_terms = (
            nw.from_native(sc_concat.df_estimates)
            .lazy()
            .select(self.variable_ids)
            .unique()
        )
        overlap = self_terms.join(
            other_terms, on=self.variable_ids, how="inner"
        ).collect()
        if overlap.shape[0] > 0:
            raise ValueError(
                "AdapterStats.concat_with(how='vertical'): self and "
                "sc_concat share at least one variable_ids value - "
                "stacking their df_vcov block-diagonally would be "
                "ambiguous for those shared terms (which side's "
                "covariance would apply?). Rename the overlapping terms "
                "on one side first if you need to keep both."
            )

        logger.warning(
            "AdapterStats.concat_with(how='vertical'): stacking df_vcov "
            "from two objects block-diagonally - self's and sc_concat's "
            "terms are assumed uncorrelated (no cross-covariance is known "
            "or represented), so a joint SE computed afterward across a "
            "self term and a sc_concat term will be wrong; joint SEs "
            "within either original object's own terms remain correct."
        )
        new_vcov = concat_wrapper([self_vcov, other_vcov], how="diagonal")

        result = super().concat_with(sc_concat, how=how)
        result.replicate_stats.df_vcov = new_vcov
        return result

    def _raise_if_vcov_or_tidy(self, method_name: str) -> None:
        present = [
            name
            for name, value in (
                ("df_vcov", self.replicate_stats.df_vcov),
                ("df_tidy", self.replicate_stats.df_tidy),
            )
            if value is not None
        ]
        if not present:
            return

        joined = " and ".join(present)
        raise ValueError(
            f"AdapterStats.{method_name}() can't reshape {joined} generically "
            f"(there's no safe, generic way to know whether/how it should "
            f"follow along) and won't silently drop it - clear it first (e.g. "
            f"self.replicate_stats.df_vcov = None) if you don't need it, or "
            f"use filter()/rename() instead, which keep df_vcov in sync."
        )

Pure Python (statsmodels, linearmodels, pyfixest, polars_ds)

No R/Stata/rpy2/pystata needed - these four wrap packages already available in Python.

statsmodels_adapter

statsmodels_adapter(
    df,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_constant: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "HC3",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
) -> AdapterStats

Fit an OLS/WLS regression with statsmodels and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
y dependent variable column name.
required
x predictor column name(s).
required
weight column name for weighted least squares (WLS), or None for

unweighted OLS. Default is None.

None
add_constant whether to add an intercept term (named "const"). Default

is True.

True
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE/covariance value column in the

output. Default is "estimate".

'estimate'
cov_type passed to `.fit()`. Defaults to "HC3" (heteroskedasticity-

robust) rather than statsmodels' own classical default, since HC3 performs well even in small-to-moderate samples and there's rarely a reason to assume homoskedasticity for survey/implicate data. Pass "nonrobust" to opt back into classical SEs.

'HC3'
cov_kwds extra keywords for the covariance estimator (e.g.

{"groups": df["cluster"]} for cluster-robust SEs). Default is None.

None
model_kwargs extra keywords forwarded to the model constructor

(sm.OLS/sm.WLS). Default is None.

None
fit_kwargs extra keywords forwarded to `.fit()` besides cov_type/

cov_kwds (e.g. {"maxiter": 200}). Default is None.

None

Returns:

Type Description
AdapterStats

(df_estimates, df_ses, df_vcov, df_tidy) - df_vcov is always populated here since statsmodels computes it for free alongside .bse. df_tidy is statsmodels' own coefficient table (results.summary2().tables[1]: Coef./Std.Err./z or t/P>|z|/CI), held as-is - a single implicate's own t-stats/p-values aren't valid MI inference, so this is a diagnostic snapshot only, never combined across implicates.

Source code in src/survey_kit/statistics/adapters.py
def statsmodels_adapter(
    df,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_constant: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "HC3",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
) -> AdapterStats:
    """
    Fit an OLS/WLS regression with statsmodels and return its coefficient
    table in survey_kit's normalized (df_estimates, df_ses, df_vcov,
    df_tidy) shape.

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    y : dependent variable column name.
    x : predictor column name(s).
    weight : column name for weighted least squares (WLS), or None for
        unweighted OLS. Default is None.
    add_constant : whether to add an intercept term (named "const"). Default
        is True.
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE/covariance value column in the
        output. Default is "estimate".
    cov_type : passed to `.fit()`. Defaults to "HC3" (heteroskedasticity-
        robust) rather than statsmodels' own classical default, since HC3
        performs well even in small-to-moderate samples and there's rarely a
        reason to assume homoskedasticity for survey/implicate data. Pass
        "nonrobust" to opt back into classical SEs.
    cov_kwds : extra keywords for the covariance estimator (e.g.
        `{"groups": df["cluster"]}` for cluster-robust SEs). Default is None.
    model_kwargs : extra keywords forwarded to the model constructor
        (`sm.OLS`/`sm.WLS`). Default is None.
    fit_kwargs : extra keywords forwarded to `.fit()` besides cov_type/
        cov_kwds (e.g. `{"maxiter": 200}`). Default is None.

    Returns
    -------
    AdapterStats
        (df_estimates, df_ses, df_vcov, df_tidy) - df_vcov is always
        populated here since statsmodels computes it for free alongside
        .bse. df_tidy is statsmodels' own coefficient table
        (`results.summary2().tables[1]`: Coef./Std.Err./z or t/P>|z|/CI),
        held as-is - a single implicate's own t-stats/p-values aren't valid
        MI inference, so this is a diagnostic snapshot only, never combined
        across implicates.
    """
    try:
        import statsmodels.api as sm
    except ImportError as e:
        message = (
            "statsmodels_adapter requires the 'statsmodels' package - "
            "install it with `uv add statsmodels`."
        )
        logger.error(message)
        raise ImportError(message) from e

    df_pd = _to_pandas(df)
    x = [x] if isinstance(x, str) else list(x)
    model_kwargs = dict(model_kwargs or {})
    fit_kwargs = dict(fit_kwargs or {})

    exog = df_pd[x]
    if add_constant:
        exog = sm.add_constant(exog, has_constant="add")

    if weight:
        model = sm.WLS(df_pd[y], exog, weights=df_pd[weight], **model_kwargs)
    else:
        model = sm.OLS(df_pd[y], exog, **model_kwargs)

    results = model.fit(cov_type=cov_type, cov_kwds=cov_kwds, **fit_kwargs)

    df_estimates, df_ses = _coef_table_from_series(
        results.params, results.bse, join_on_name, value_name
    )
    df_vcov = _vcov_table_from_frame(results.cov_params(), join_on_name, value_name)
    df_tidy = pl.from_pandas(
        results.summary2().tables[1].reset_index(names=join_on_name)
    )

    return _adapter_stats(df_estimates, df_ses, df_vcov, df_tidy, join_on_name)

mi_ses_from_statsmodels

mi_ses_from_statsmodels(
    df_implicates,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_constant: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "HC3",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
)

mi_ses_from_function(delegate=statsmodels_adapter, ...), with statsmodels_adapter's own arguments taken directly as keyword arguments instead of packed into an arguments={} dict.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
y str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
x str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
weight str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
add_constant str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
join_on_name str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
value_name str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
cov_type str

cov_kwds, model_kwargs, fit_kwargs : see statsmodels_adapter.

required
replicates `survey_kit.statistics.replicates.Replicates`, optional.

When given, per-implicate SEs come from resampling across replicate weights instead of statsmodels' own cov_type/cov_kwds - statsmodels_adapter is run once per replicate weight column (point estimates only), the same replicate-weight-bootstrap approach mi_ses_from_stata's replicates= uses. weight is ignored in this mode (each replicate call supplies its own weight column) - the implicate is converted to pandas once, not once per replicate. Default is None (use statsmodels_adapter's own cov_type/cov_kwds).

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

statsmodels_adapter : the delegate this wraps.

Source code in src/survey_kit/statistics/adapters.py
def mi_ses_from_statsmodels(
    df_implicates,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_constant: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "HC3",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
):
    """
    `mi_ses_from_function(delegate=statsmodels_adapter, ...)`, with
    `statsmodels_adapter`'s own arguments taken directly as keyword
    arguments instead of packed into an `arguments={}` dict.

    Parameters
    ----------
    df_implicates, path_srmi, index, df_noimputes, parallel,
        parallel_inputs, rounding, round_output : see
        [`mi_ses_from_function`][survey_kit.statistics.multiple_imputation.mi_ses_from_function].
    y, x, weight, add_constant, join_on_name, value_name, cov_type,
        cov_kwds, model_kwargs, fit_kwargs : see `statsmodels_adapter`.
    replicates : `survey_kit.statistics.replicates.Replicates`, optional.
        When given, per-implicate SEs come from resampling across
        replicate weights instead of statsmodels' own cov_type/cov_kwds -
        `statsmodels_adapter` is run once per replicate weight column
        (point estimates only), the same replicate-weight-bootstrap
        approach `mi_ses_from_stata`'s `replicates=` uses. `weight` is
        ignored in this mode (each replicate call supplies its own weight
        column) - the implicate is converted to pandas once, not once per
        replicate. Default is None (use `statsmodels_adapter`'s own
        cov_type/cov_kwds).

    Returns
    -------
    MultipleImputation
        Same as `mi_ses_from_function`.

    See Also
    --------
    statsmodels_adapter : the delegate this wraps.
    """
    from .multiple_imputation import mi_ses_from_function

    base_arguments = {
        "y": y,
        "x": x,
        "add_constant": add_constant,
        "join_on_name": join_on_name,
        "value_name": value_name,
        "model_kwargs": model_kwargs,
        "fit_kwargs": fit_kwargs,
    }

    if replicates is None:
        delegate = statsmodels_adapter
        arguments = {
            **base_arguments,
            "weight": weight,
            "cov_type": cov_type,
            "cov_kwds": cov_kwds,
        }
    else:
        if weight:
            logger.warning(
                "mi_ses_from_statsmodels: weight is ignored when replicates "
                "is given - each replicate call supplies its own weight column."
            )
        #   cov_type/cov_kwds are discarded for these point-estimate-only
        #   replicate calls (the SE comes from the spread across
        #   replicates, not statsmodels' own covariance) - force the
        #   cheapest (classical) option rather than whatever the caller
        #   passed, to skip computing a robust covariance matrix that's
        #   about to be thrown away, on every one of potentially dozens of
        #   replicate calls.
        delegate = _mi_ses_replicates_delegate(
            statsmodels_adapter,
            {**base_arguments, "cov_type": "nonrobust", "cov_kwds": None},
            join_on_name,
            replicates,
            convert=_to_pandas,
        )
        arguments = {}

    return mi_ses_from_function(
        delegate=delegate,
        df_implicates=df_implicates,
        join_on=[join_on_name],
        path_srmi=path_srmi,
        index=index,
        df_noimputes=df_noimputes,
        arguments=arguments,
        parallel=parallel,
        parallel_inputs=parallel_inputs,
        rounding=rounding,
        round_output=round_output,
    )

linearmodels_adapter

linearmodels_adapter(
    df,
    formula: str,
    weight: str | None = None,
    model: str = "IV2SLS",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "robust",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
) -> AdapterStats

Fit a linearmodels model (IV/panel) and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.

Unlike statsmodels_adapter/polars_ds_adapter, this takes a formula string rather than y/x lists - linearmodels' bracket syntax ("y ~ 1 + x1 + [x2 ~ z1 + z2]" for instrumenting x2 with z1/z2) is how IV/panel specifications are expressed, and there's no y/x-list equivalent for that.

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
formula linearmodels formula string, e.g. "y ~ 1 + x1 + x2" for plain

OLS via IV2SLS, or with a bracketed [endog ~ instruments] clause for actual instrumental variables.

required
weight column name for weighted estimation, or None. Default is None.
None
model one of "IV2SLS", "PanelOLS", "PooledOLS", "RandomEffects",

"BetweenOLS". Default is "IV2SLS" (also covers plain OLS, with no instruments in the formula). Panel models require df to have the entity/time MultiIndex they expect - set that up before calling mi_ses_from_function (e.g. via a pre/index-setting step), this adapter doesn't build one for you.

'IV2SLS'
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE/covariance value column in the

output. Default is "estimate".

'estimate'
cov_type passed to `.fit()`. Defaults to "robust" - linearmodels' own

vocabulary for heteroskedasticity-consistent SEs ("HC0-3" is a statsmodels-specific term with no direct equivalent here); "robust" is the closest counterpart to statsmodels_adapter's HC3 default, for the same reason (rarely safe to assume homoskedasticity).

'robust'
cov_kwds extra keywords for the covariance estimator (e.g.

{"clusters": df["cluster"]} for cluster-robust SEs), spread as **cov_kwds into .fit() since linearmodels takes them as loose kwargs rather than a nested dict. Default is None.

None
model_kwargs extra keywords forwarded to the model constructor.

Default is None.

None
fit_kwargs extra keywords forwarded to `.fit()` besides cov_type/

cov_kwds. Default is None.

None

Returns:

Type Description
AdapterStats

(df_estimates, df_ses, df_vcov, df_tidy) - df_vcov is always populated here since linearmodels computes it for free alongside .std_errors. df_tidy assembles estimate/std_error/statistic/ p_value/conf_low/conf_high from linearmodels' own .params/ .std_errors/.tstats/.pvalues/.conf_int() (linearmodels has no single ready-made tidy table the way statsmodels/pyfixest do) - a diagnostic snapshot only, never combined across implicates.

Source code in src/survey_kit/statistics/adapters.py
def linearmodels_adapter(
    df,
    formula: str,
    weight: str | None = None,
    model: str = "IV2SLS",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "robust",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
) -> AdapterStats:
    """
    Fit a linearmodels model (IV/panel) and return its coefficient table in
    survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.

    Unlike statsmodels_adapter/polars_ds_adapter, this takes a formula
    string rather than y/x lists - linearmodels' bracket syntax
    (`"y ~ 1 + x1 + [x2 ~ z1 + z2]"` for instrumenting x2 with z1/z2) is how
    IV/panel specifications are expressed, and there's no y/x-list
    equivalent for that.

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    formula : linearmodels formula string, e.g. "y ~ 1 + x1 + x2" for plain
        OLS via IV2SLS, or with a bracketed `[endog ~ instruments]` clause
        for actual instrumental variables.
    weight : column name for weighted estimation, or None. Default is None.
    model : one of "IV2SLS", "PanelOLS", "PooledOLS", "RandomEffects",
        "BetweenOLS". Default is "IV2SLS" (also covers plain OLS, with no
        instruments in the formula). Panel models require df to have the
        entity/time MultiIndex they expect - set that up before calling
        mi_ses_from_function (e.g. via a `pre`/index-setting step), this
        adapter doesn't build one for you.
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE/covariance value column in the
        output. Default is "estimate".
    cov_type : passed to `.fit()`. Defaults to "robust" - linearmodels' own
        vocabulary for heteroskedasticity-consistent SEs ("HC0-3" is a
        statsmodels-specific term with no direct equivalent here); "robust"
        is the closest counterpart to statsmodels_adapter's HC3 default,
        for the same reason (rarely safe to assume homoskedasticity).
    cov_kwds : extra keywords for the covariance estimator (e.g.
        `{"clusters": df["cluster"]}` for cluster-robust SEs), spread as
        `**cov_kwds` into `.fit()` since linearmodels takes them as loose
        kwargs rather than a nested dict. Default is None.
    model_kwargs : extra keywords forwarded to the model constructor.
        Default is None.
    fit_kwargs : extra keywords forwarded to `.fit()` besides cov_type/
        cov_kwds. Default is None.

    Returns
    -------
    AdapterStats
        (df_estimates, df_ses, df_vcov, df_tidy) - df_vcov is always
        populated here since linearmodels computes it for free alongside
        .std_errors. df_tidy assembles estimate/std_error/statistic/
        p_value/conf_low/conf_high from linearmodels' own .params/
        .std_errors/.tstats/.pvalues/.conf_int() (linearmodels has no
        single ready-made tidy table the way statsmodels/pyfixest do) - a
        diagnostic snapshot only, never combined across implicates.
    """
    try:
        from linearmodels.iv import IV2SLS
        from linearmodels.panel import BetweenOLS, PanelOLS, PooledOLS, RandomEffects
    except ImportError as e:
        message = (
            "linearmodels_adapter requires the 'linearmodels' package - "
            "install it with `uv add linearmodels`."
        )
        logger.error(message)
        raise ImportError(message) from e

    model_classes = {
        "IV2SLS": IV2SLS,
        "PanelOLS": PanelOLS,
        "PooledOLS": PooledOLS,
        "RandomEffects": RandomEffects,
        "BetweenOLS": BetweenOLS,
    }
    if model not in model_classes:
        message = f"Unknown linearmodels model '{model}'; expected one of {list(model_classes)}."
        logger.error(message)
        raise ValueError(message)

    df_pd = _to_pandas(df)
    model_kwargs = dict(model_kwargs or {})
    fit_kwargs = dict(fit_kwargs or {})
    cov_kwds = dict(cov_kwds or {})

    if weight:
        model_kwargs["weights"] = df_pd[weight]

    model_obj = model_classes[model].from_formula(formula, data=df_pd, **model_kwargs)
    results = model_obj.fit(cov_type=cov_type, **cov_kwds, **fit_kwargs)

    df_estimates, df_ses = _coef_table_from_series(
        results.params, results.std_errors, join_on_name, value_name
    )
    df_vcov = _vcov_table_from_frame(results.cov, join_on_name, value_name)
    df_tidy = _tidy_from_parts(
        join_on_name,
        results.params,
        results.std_errors,
        results.tstats,
        results.pvalues,
        results.conf_int(),
    )

    return _adapter_stats(df_estimates, df_ses, df_vcov, df_tidy, join_on_name)

mi_ses_from_linearmodels

mi_ses_from_linearmodels(
    df_implicates,
    formula: str,
    weight: str | None = None,
    model: str = "IV2SLS",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "robust",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
)

mi_ses_from_function(delegate=linearmodels_adapter, ...), with linearmodels_adapter's own arguments taken directly as keyword arguments instead of packed into an arguments={} dict.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
formula str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
weight str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
model str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
join_on_name str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
value_name str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
cov_type str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
cov_kwds str

model_kwargs, fit_kwargs : see linearmodels_adapter.

required
replicates `survey_kit.statistics.replicates.Replicates`, optional.

When given, per-implicate SEs come from resampling across replicate weights instead of linearmodels' own cov_type/cov_kwds - linearmodels_adapter is run once per replicate weight column (point estimates only, with cov_type forced to "unadjusted" since the covariance is discarded anyway), the same replicate-weight- bootstrap approach mi_ses_from_stata's replicates= uses. weight is ignored in this mode. Default is None (use linearmodels_adapter's own cov_type/cov_kwds).

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

linearmodels_adapter : the delegate this wraps.

Source code in src/survey_kit/statistics/adapters.py
def mi_ses_from_linearmodels(
    df_implicates,
    formula: str,
    weight: str | None = None,
    model: str = "IV2SLS",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    cov_type: str = "robust",
    cov_kwds: dict | None = None,
    model_kwargs: dict | None = None,
    fit_kwargs: dict | None = None,
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
):
    """
    `mi_ses_from_function(delegate=linearmodels_adapter, ...)`, with
    `linearmodels_adapter`'s own arguments taken directly as keyword
    arguments instead of packed into an `arguments={}` dict.

    Parameters
    ----------
    df_implicates, path_srmi, index, df_noimputes, parallel,
        parallel_inputs, rounding, round_output : see
        [`mi_ses_from_function`][survey_kit.statistics.multiple_imputation.mi_ses_from_function].
    formula, weight, model, join_on_name, value_name, cov_type, cov_kwds,
        model_kwargs, fit_kwargs : see `linearmodels_adapter`.
    replicates : `survey_kit.statistics.replicates.Replicates`, optional.
        When given, per-implicate SEs come from resampling across
        replicate weights instead of linearmodels' own cov_type/cov_kwds -
        `linearmodels_adapter` is run once per replicate weight column
        (point estimates only, with cov_type forced to "unadjusted" since
        the covariance is discarded anyway), the same replicate-weight-
        bootstrap approach `mi_ses_from_stata`'s `replicates=` uses.
        `weight` is ignored in this mode. Default is None (use
        `linearmodels_adapter`'s own cov_type/cov_kwds).

    Returns
    -------
    MultipleImputation
        Same as `mi_ses_from_function`.

    See Also
    --------
    linearmodels_adapter : the delegate this wraps.
    """
    from .multiple_imputation import mi_ses_from_function

    base_arguments = {
        "formula": formula,
        "model": model,
        "join_on_name": join_on_name,
        "value_name": value_name,
        "model_kwargs": model_kwargs,
        "fit_kwargs": fit_kwargs,
    }

    if replicates is None:
        delegate = linearmodels_adapter
        arguments = {
            **base_arguments,
            "weight": weight,
            "cov_type": cov_type,
            "cov_kwds": cov_kwds,
        }
    else:
        if weight:
            logger.warning(
                "mi_ses_from_linearmodels: weight is ignored when replicates "
                "is given - each replicate call supplies its own weight column."
            )
        delegate = _mi_ses_replicates_delegate(
            linearmodels_adapter,
            {**base_arguments, "cov_type": "unadjusted", "cov_kwds": None},
            join_on_name,
            replicates,
            convert=_to_pandas,
        )
        arguments = {}

    return mi_ses_from_function(
        delegate=delegate,
        df_implicates=df_implicates,
        join_on=[join_on_name],
        path_srmi=path_srmi,
        index=index,
        df_noimputes=df_noimputes,
        arguments=arguments,
        parallel=parallel,
        parallel_inputs=parallel_inputs,
        rounding=rounding,
        round_output=round_output,
    )

pyfixest_adapter

pyfixest_adapter(
    df,
    formula: str,
    func: str = "feols",
    family: str | None = None,
    weight: str | None = None,
    vcov: str | dict | None = "hetero",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **kwargs,
) -> AdapterStats

Fit a pyfixest regression and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape. pyfixest mirrors R's fixest syntax/functionality (fixed effects via formula, e.g. "y ~ x1 | firm", robust/clustered SEs, OLS/GLM/Poisson) natively in Python - no R/rpy2 needed, and generally the better default over r_feols/r_feglm/r_fepois/r_fixest_adapter unless you specifically need something pyfixest doesn't cover (those pull in a full R + rpy2 + rpy2-arrow dependency chain for no extra benefit if pyfixest already does the job).

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
formula fixest-syntax formula string, e.g. "y ~ x1 + x2 | firm" for a

fit with a firm fixed effect.

required
func which pyfixest estimator to call - "feols" (default), "feglm",

or "fepois". Called as getattr(pyfixest, func)(...).

'feols'
family passed to feglm as its required `family=` argument (e.g.

"logit", "probit", "poisson"). Not used for feols/fepois - leave as None.

None
weight column name for weighted estimation, or None. Passed straight

through as pyfixest's own weights= (a plain column name - no formula/raw-string gymnastics needed, unlike the R adapters).

None
vcov pyfixest's own `vcov=` argument - a string like "hetero" (robust,

the default here for the same reason statsmodels_adapter defaults to HC3: rarely safe to assume homoskedasticity), "iid" (classical), or a dict for clustering, e.g. {"CRV1": "firm"}.

'hetero'
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE/covariance value column in the

output. Default is "estimate".

'estimate'
**kwargs any other argument the chosen estimator takes (`ssc`,

fixef_rm, split/fsplit, offset for fepois, ...) - forwarded verbatim, no conversion needed since this stays in pure Python.

{}

Returns:

Type Description
AdapterStats

(df_estimates, df_ses, df_vcov, df_tidy). df_vcov is populated from the fitted model's internal covariance matrix when its shape matches the coefficient vector, else None (with a warning) rather than risking a silent term misalignment - pyfixest doesn't expose this as public API, so this reads a private attribute defensively. df_tidy is pyfixest's own fit.tidy() table (Estimate/Std. Error/ t value/Pr(>|t|)/CI) - a diagnostic snapshot only, never combined across implicates.

Source code in src/survey_kit/statistics/adapters.py
def pyfixest_adapter(
    df,
    formula: str,
    func: str = "feols",
    family: str | None = None,
    weight: str | None = None,
    vcov: str | dict | None = "hetero",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **kwargs,
) -> AdapterStats:
    """
    Fit a pyfixest regression and return its coefficient table in
    survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.
    pyfixest mirrors R's fixest syntax/functionality (fixed effects via formula,
    e.g. "y ~ x1 | firm", robust/clustered SEs, OLS/GLM/Poisson) natively in
    Python - no R/rpy2 needed, and generally the better default over
    `r_feols`/`r_feglm`/`r_fepois`/`r_fixest_adapter` unless you specifically
    need something pyfixest doesn't cover (those pull in a full R + rpy2 +
    rpy2-arrow dependency chain for no extra benefit if pyfixest already
    does the job).

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    formula : fixest-syntax formula string, e.g. "y ~ x1 + x2 | firm" for a
        fit with a firm fixed effect.
    func : which pyfixest estimator to call - "feols" (default), "feglm",
        or "fepois". Called as `getattr(pyfixest, func)(...)`.
    family : passed to feglm as its required `family=` argument (e.g.
        "logit", "probit", "poisson"). Not used for feols/fepois - leave as
        None.
    weight : column name for weighted estimation, or None. Passed straight
        through as pyfixest's own `weights=` (a plain column name - no
        formula/raw-string gymnastics needed, unlike the R adapters).
    vcov : pyfixest's own `vcov=` argument - a string like "hetero" (robust,
        the default here for the same reason statsmodels_adapter defaults
        to HC3: rarely safe to assume homoskedasticity), "iid" (classical),
        or a dict for clustering, e.g. `{"CRV1": "firm"}`.
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE/covariance value column in the
        output. Default is "estimate".
    **kwargs : any other argument the chosen estimator takes (`ssc`,
        `fixef_rm`, `split`/`fsplit`, `offset` for fepois, ...) - forwarded
        verbatim, no conversion needed since this stays in pure Python.

    Returns
    -------
    AdapterStats
        (df_estimates, df_ses, df_vcov, df_tidy). df_vcov is populated from
        the fitted model's internal covariance matrix when its shape
        matches the coefficient vector, else None (with a warning) rather
        than risking a silent term misalignment - pyfixest doesn't expose
        this as public API, so this reads a private attribute defensively.
        df_tidy is pyfixest's own `fit.tidy()` table (Estimate/Std. Error/
        t value/Pr(>|t|)/CI) - a diagnostic snapshot only, never combined
        across implicates.
    """
    try:
        import pyfixest as pf
    except ImportError as e:
        message = (
            "pyfixest_adapter requires the 'pyfixest' package - install it "
            "with `uv add pyfixest` (or `pip install survey-kit[pyfixest]`)."
        )
        logger.error(message)
        raise ImportError(message) from e

    df_pd = _to_pandas(df)

    call_kwargs = dict(kwargs)
    if weight is not None:
        call_kwargs["weights"] = weight
    if family is not None:
        call_kwargs["family"] = family
    if vcov is not None:
        call_kwargs["vcov"] = vcov

    fit = getattr(pf, func)(formula, data=df_pd, **call_kwargs)

    coef = fit.coef()
    se = fit.se()

    df_estimates = pl.DataFrame(
        {join_on_name: list(coef.index), value_name: coef.to_numpy()}
    )
    df_ses = pl.DataFrame({join_on_name: list(se.index), value_name: se.to_numpy()})

    df_vcov = None
    vcov_matrix = getattr(fit, "_vcov", None)
    if vcov_matrix is not None:
        import numpy as np

        terms = list(coef.index)
        n = len(terms)
        vcov_matrix = np.asarray(vcov_matrix)
        if vcov_matrix.shape == (n, n):
            df_vcov = pl.DataFrame(
                [
                    {
                        f"{join_on_name}_1": terms[i],
                        f"{join_on_name}_2": terms[j],
                        value_name: float(vcov_matrix[i, j]),
                    }
                    for i in range(n)
                    for j in range(n)
                ]
            )
        else:
            logger.warning(
                "pyfixest_adapter: the fitted model's internal covariance "
                f"matrix shape {vcov_matrix.shape} doesn't match the "
                f"{n} coefficient terms - returning df_vcov=None rather "
                "than risk misaligning terms."
            )

    df_tidy = pl.from_pandas(fit.tidy().reset_index(names=join_on_name))

    return _adapter_stats(df_estimates, df_ses, df_vcov, df_tidy, join_on_name)

mi_ses_from_pyfixest

Namespace for per-estimator mi_ses_from_function(delegate= pyfixest_adapter, ...) wrappers - mi_ses_from_pyfixest.feols(...), .fepois(...), .feglm(...) - each naming the common pyfixest_adapter arguments (fml, weight, vcov, family) explicitly instead of picking the estimator via a func="..." string, so IDEs show the right parameters for the one you're actually calling. Everything else pyfixest.feols/fepois/feglm accepts still flows through **kwargs exactly as in pyfixest_adapter.

feglm staticmethod

feglm(
    df_implicates,
    fml: str,
    family: str,
    vcov: str | dict | None = "hetero",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **kwargs,
)

mi_ses_from_function(delegate=pyfixest_adapter, ...) for pyfixest.feglm specifically - GLM with fixed effects (family= required, e.g. "logit", "probit", "poisson"). No replicates= option here (unlike .feols/.fepois) - pyfixest's feglm() has no weights= argument to substitute a replicate weight column into.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
fml see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
family see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
vcov see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
join_on_name see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
value_name see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
**kwargs any other `pyfixest.feglm` argument (`ssc`,

fixef_rm, iwls_tol/iwls_maxiter, ...) - forwarded verbatim, see pyfixest_adapter.

{}

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

pyfixest_adapter : the delegate this wraps.

feols staticmethod

feols(
    df_implicates,
    fml: str,
    weight: str | None = None,
    vcov: str | dict | None = "hetero",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **kwargs,
)

mi_ses_from_function(delegate=pyfixest_adapter, ...) for pyfixest.feols specifically - OLS/IV with fixed effects, fixest formula syntax (e.g. "y ~ x1 + x2 | firm").

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
fml see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
weight see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
vcov see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
join_on_name see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
value_name see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
replicates `survey_kit.statistics.replicates.Replicates`,

optional. When given, per-implicate SEs come from resampling across replicate weights instead of vcov - pyfixest_adapter is run once per replicate weight column (point estimates only, with vcov forced to "iid" since it's discarded anyway), the same replicate-weight-bootstrap approach mi_ses_from_stata's replicates= uses. weight is ignored in this mode. Default is None (use pyfixest_adapter's own vcov).

None
**kwargs any other `pyfixest.feols` argument (`ssc`, `fixef_rm`,

split/fsplit, ...) - forwarded verbatim, see pyfixest_adapter.

{}

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

pyfixest_adapter : the delegate this wraps.

fepois staticmethod

fepois(
    df_implicates,
    fml: str,
    weight: str | None = None,
    vcov: str | dict | None = "hetero",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **kwargs,
)

mi_ses_from_function(delegate=pyfixest_adapter, ...) for pyfixest.fepois specifically - Poisson regression with fixed effects.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
fml see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
weight see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
vcov see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
join_on_name see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
value_name see

pyfixest_adapter (fml is pyfixest_adapter's formula).

required
replicates see `mi_ses_from_pyfixest.feols` - same

replicate-weight-bootstrap option, vcov forced to "iid" and weight ignored in this mode. Default is None.

None
**kwargs any other `pyfixest.fepois` argument (`ssc`,

fixef_rm, offset, iwls_tol/iwls_maxiter, ...) - forwarded verbatim, see pyfixest_adapter.

{}

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

pyfixest_adapter : the delegate this wraps.

polars_ds_adapter

polars_ds_adapter(
    df,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_bias: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    std_err: str = "hc3",
    null_policy: str = "raise",
) -> AdapterStats

Fit an OLS/WLS regression with polars_ds's lin_reg_report and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, None, df_tidy) shape. Stays entirely in polars/narwhals - no pandas conversion.

polars_ds doesn't expose a coefficient covariance matrix, so df_vcov is always None here: contrasts between two terms of the same fit (.compare(other, compare_list_variables=[...])) aren't calculable from this adapter's output alone. Use statsmodels_adapter or linearmodels_adapter if you need that.

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
y dependent variable column name.
required
x predictor column name(s).
required
weight column name for weighted least squares, or None. **Note**: when

weight is given, polars_ds always falls back to homoskedastic standard errors regardless of std_err - a limitation of the underlying library (it doesn't yet implement weighted HC0-3), not of this adapter. A warning is logged when this happens.

None
add_bias whether to add an intercept term (named "__bias__", polars_ds's

own naming). Default is True.

True
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE value column in the output.

Default is "estimate".

'estimate'
std_err one of "se" (classical/homoskedastic), "hc0", "hc1", "hc2",

"hc3". Defaults to "hc3", for the same reason statsmodels_adapter defaults there - rarely safe to assume homoskedasticity. Silently ignored (falls back to "se") when weight is given - see the weight parameter above.

'hc3'
null_policy how to handle nulls in the predictors, passed straight to

polars_ds. Default is "raise".

'raise'

Returns:

Type Description
AdapterStats

df_vcov is always None here - df_tidy is polars_ds's full lin_reg_report output as-is (features/beta/se/t/p/CI/r2/adj_r2), just with its term column renamed to join_on_name - a diagnostic snapshot only, never combined across implicates.

Source code in src/survey_kit/statistics/adapters.py
def polars_ds_adapter(
    df,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_bias: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    std_err: str = "hc3",
    null_policy: str = "raise",
) -> AdapterStats:
    """
    Fit an OLS/WLS regression with polars_ds's `lin_reg_report` and return
    its coefficient table in survey_kit's normalized (df_estimates, df_ses,
    None, df_tidy) shape. Stays entirely in polars/narwhals - no pandas
    conversion.

    polars_ds doesn't expose a coefficient covariance matrix, so df_vcov is
    always None here: contrasts between two terms of the same fit
    (`.compare(other, compare_list_variables=[...])`) aren't calculable from
    this adapter's output alone. Use statsmodels_adapter or
    linearmodels_adapter if you need that.

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    y : dependent variable column name.
    x : predictor column name(s).
    weight : column name for weighted least squares, or None. **Note**: when
        weight is given, polars_ds always falls back to homoskedastic
        standard errors regardless of `std_err` - a limitation of the
        underlying library (it doesn't yet implement weighted HC0-3), not
        of this adapter. A warning is logged when this happens.
    add_bias : whether to add an intercept term (named "__bias__", polars_ds's
        own naming). Default is True.
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE value column in the output.
        Default is "estimate".
    std_err : one of "se" (classical/homoskedastic), "hc0", "hc1", "hc2",
        "hc3". Defaults to "hc3", for the same reason statsmodels_adapter
        defaults there - rarely safe to assume homoskedasticity. Silently
        ignored (falls back to "se") when weight is given - see the weight
        parameter above.
    null_policy : how to handle nulls in the predictors, passed straight to
        polars_ds. Default is "raise".

    Returns
    -------
    AdapterStats
        df_vcov is always None here - df_tidy is polars_ds's full
        `lin_reg_report` output as-is (features/beta/se/t/p/CI/r2/adj_r2),
        just with its term column renamed to join_on_name - a diagnostic
        snapshot only, never combined across implicates.
    """
    try:
        import polars_ds as pds
    except ImportError as e:
        message = (
            "polars_ds_adapter requires the 'polars_ds' package - "
            "install it with `uv add polars_ds`."
        )
        logger.error(message)
        raise ImportError(message) from e

    if weight and std_err != "se":
        logger.warning(
            "polars_ds ignores std_err when weight is given and falls back to "
            "homoskedastic standard errors - see polars_ds_adapter's docstring."
        )

    x = [x] if isinstance(x, str) else list(x)

    if isinstance(df, pl.DataFrame):
        df_pl = df.lazy()
    elif isinstance(df, pl.LazyFrame):
        df_pl = df
    else:
        import narwhals as nw

        df_pl = nw.from_native(df).lazy().collect().to_polars().lazy()

    report = (
        df_pl.select(
            pds.lin_reg_report(
                *x,
                target=y,
                weights=weight,
                add_bias=add_bias,
                null_policy=null_policy,
                std_err=std_err,
            ).alias("__report__")
        )
        .collect()
        .unnest("__report__")
    )

    se_col = "std_err" if (weight or std_err == "se") else f"{std_err}_se"

    df_estimates = report.select(
        pl.col("features").alias(join_on_name), pl.col("beta").alias(value_name)
    )
    df_ses = report.select(
        pl.col("features").alias(join_on_name), pl.col(se_col).alias(value_name)
    )
    df_tidy = report.rename({"features": join_on_name})

    return _adapter_stats(df_estimates, df_ses, None, df_tidy, join_on_name)

mi_ses_from_polars_ds

mi_ses_from_polars_ds(
    df_implicates,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_bias: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    std_err: str = "hc3",
    null_policy: str = "raise",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
)

mi_ses_from_function(delegate=polars_ds_adapter, ...), with polars_ds_adapter's own arguments taken directly as keyword arguments instead of packed into an arguments={} dict.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
y str

see polars_ds_adapter.

required
x str

see polars_ds_adapter.

required
weight str

see polars_ds_adapter.

required
add_bias str

see polars_ds_adapter.

required
join_on_name str

see polars_ds_adapter.

required
value_name str

see polars_ds_adapter.

required
std_err str

see polars_ds_adapter.

required
null_policy str

see polars_ds_adapter.

required
replicates `survey_kit.statistics.replicates.Replicates`, optional.

When given, per-implicate SEs come from resampling across replicate weights instead of std_err - polars_ds_adapter is run once per replicate weight column (point estimates only, with std_err forced to "se" since it's discarded anyway), the same replicate-weight-bootstrap approach mi_ses_from_stata's replicates= uses. weight is ignored in this mode; the implicate is converted to polars once (if it isn't already), not once per replicate. Default is None (use polars_ds_adapter's own std_err).

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

polars_ds_adapter : the delegate this wraps.

Source code in src/survey_kit/statistics/adapters.py
def mi_ses_from_polars_ds(
    df_implicates,
    y: str,
    x: list[str] | str,
    weight: str | None = None,
    add_bias: bool = True,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    std_err: str = "hc3",
    null_policy: str = "raise",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
):
    """
    `mi_ses_from_function(delegate=polars_ds_adapter, ...)`, with
    `polars_ds_adapter`'s own arguments taken directly as keyword
    arguments instead of packed into an `arguments={}` dict.

    Parameters
    ----------
    df_implicates, path_srmi, index, df_noimputes, parallel,
        parallel_inputs, rounding, round_output : see
        [`mi_ses_from_function`][survey_kit.statistics.multiple_imputation.mi_ses_from_function].
    y, x, weight, add_bias, join_on_name, value_name, std_err, null_policy :
        see `polars_ds_adapter`.
    replicates : `survey_kit.statistics.replicates.Replicates`, optional.
        When given, per-implicate SEs come from resampling across
        replicate weights instead of `std_err` - `polars_ds_adapter` is
        run once per replicate weight column (point estimates only, with
        `std_err` forced to "se" since it's discarded anyway), the same
        replicate-weight-bootstrap approach `mi_ses_from_stata`'s
        `replicates=` uses. `weight` is ignored in this mode; the
        implicate is converted to polars once (if it isn't already), not
        once per replicate. Default is None (use `polars_ds_adapter`'s own
        `std_err`).

    Returns
    -------
    MultipleImputation
        Same as `mi_ses_from_function`.

    See Also
    --------
    polars_ds_adapter : the delegate this wraps.
    """
    from .multiple_imputation import mi_ses_from_function

    base_arguments = {
        "y": y,
        "x": x,
        "add_bias": add_bias,
        "join_on_name": join_on_name,
        "value_name": value_name,
        "null_policy": null_policy,
    }

    if replicates is None:
        delegate = polars_ds_adapter
        arguments = {**base_arguments, "weight": weight, "std_err": std_err}
    else:
        if weight:
            logger.warning(
                "mi_ses_from_polars_ds: weight is ignored when replicates "
                "is given - each replicate call supplies its own weight column."
            )
        delegate = _mi_ses_replicates_delegate(
            polars_ds_adapter,
            {**base_arguments, "std_err": "se"},
            join_on_name,
            replicates,
            convert=_to_polars,
        )
        arguments = {}

    return mi_ses_from_function(
        delegate=delegate,
        df_implicates=df_implicates,
        join_on=[join_on_name],
        path_srmi=path_srmi,
        index=index,
        df_noimputes=df_noimputes,
        arguments=arguments,
        parallel=parallel,
        parallel_inputs=parallel_inputs,
        rounding=rounding,
        round_output=round_output,
    )

R (rpy2)

Requires R itself plus rpy2/rpy2-arrow (pip install survey-kit[r]). r_feols/r_feglm/r_fepois/r_femlm (and their mi_ses_from_r_fixest shortcuts) cover fixest's four estimators with named arguments; r_fixest_adapter is the fully generic func= escape hatch for anything else fixest offers (feNmlm, feglm.fit, ...); r_lm_adapter covers base R's lm()/glm(). For any other R package/function entirely, see _r_interop and r_arbitrary_estimators.py.

r_lm_adapter

r_lm_adapter(
    df,
    formula: str,
    weight: str | None = None,
    family: str | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats

Fit a base-R lm()/glm() model and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape. Needs only R itself (no extra R packages) plus rpy2/rpy2-arrow on the Python side - see survey_kit.statistics._r_interop.check_r_setup().

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
formula R formula string, e.g. "y ~ x1 + x2".
required
weight column name for weighted least squares, or None. Default is

None. Passed as lm()/glm()'s weights= argument, referencing the column directly in R (base R's non-standard evaluation resolves it against data=, the same way the formula itself does) - unlike most other arguments here, this can't be a plain quoted string.

None
family a plain R family name (e.g. "binomial", "poisson") to fit via

glm() instead of lm() - base R resolves the string to the family function itself. For a non-default link function, wrap the full expression in _r_interop.RRaw, e.g. family=RRaw('binomial(link="probit")'). Default is None (lm(), i.e. OLS/WLS).

None
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE/covariance value column in the

output. Default is "estimate".

'estimate'
**r_kwargs any other lm()/glm() argument (e.g. `subset=`, `na.action=`,

offset=), converted to R via _r_interop.py_to_r_literal: None omits the argument, bool/int/float/str/list/dict convert naturally, a string starting with "~" is passed through raw (a formula), and _r_interop.RRaw wraps any other literal R code you need verbatim (e.g. a bare column reference). R argument names with a "." (e.g. na.action) can't be Python keyword names - pass those via a dict: r_kwargs={"na.action": ...} merged into this call, or just use **{"na.action": ...}.

{}

Returns:

Type Description
AdapterStats

(df_estimates, df_ses, df_vcov, df_tidy) - df_vcov is always populated here since vcov() is free alongside coef() in R. df_tidy is R's own summary(fit)$coefficients (Estimate/Std. Error/t or z value/Pr(>|.|)) - a diagnostic snapshot only, never combined across implicates.

Notes

Base R's lm()/glm() only provide classical (non-robust) standard errors - there's no HC0-3 equivalent without the 'sandwich' package, which this adapter deliberately doesn't pull in (keeping the R-side dependency at just R itself). Use r_fixest_adapter for built-in robust/cluster SEs.

Source code in src/survey_kit/statistics/adapters.py
def r_lm_adapter(
    df,
    formula: str,
    weight: str | None = None,
    family: str | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats:
    """
    Fit a base-R lm()/glm() model and return its coefficient table in
    survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.
    Needs only R itself (no extra R packages) plus rpy2/rpy2-arrow on the
    Python side - see `survey_kit.statistics._r_interop.check_r_setup()`.

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    formula : R formula string, e.g. "y ~ x1 + x2".
    weight : column name for weighted least squares, or None. Default is
        None. Passed as lm()/glm()'s `weights=` argument, referencing the
        column directly in R (base R's non-standard evaluation resolves it
        against `data=`, the same way the formula itself does) - unlike
        most other arguments here, this can't be a plain quoted string.
    family : a plain R family name (e.g. "binomial", "poisson") to fit via
        glm() instead of lm() - base R resolves the string to the family
        function itself. For a non-default link function, wrap the full
        expression in `_r_interop.RRaw`, e.g.
        `family=RRaw('binomial(link="probit")')`. Default is None (lm(),
        i.e. OLS/WLS).
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE/covariance value column in the
        output. Default is "estimate".
    **r_kwargs : any other lm()/glm() argument (e.g. `subset=`, `na.action=`,
        `offset=`), converted to R via `_r_interop.py_to_r_literal`:
        None omits the argument, bool/int/float/str/list/dict convert
        naturally, a string starting with "~" is passed through raw (a
        formula), and `_r_interop.RRaw` wraps any other
        literal R code you need verbatim (e.g. a bare column reference).
        R argument names with a "." (e.g. `na.action`) can't be Python
        keyword names - pass those via a dict: `r_kwargs={"na.action": ...}`
        merged into this call, or just use `**{"na.action": ...}`.

    Returns
    -------
    AdapterStats
        (df_estimates, df_ses, df_vcov, df_tidy) - df_vcov is always
        populated here since vcov() is free alongside coef() in R. df_tidy
        is R's own `summary(fit)$coefficients` (Estimate/Std. Error/t or z
        value/Pr(>|.|)) - a diagnostic snapshot only, never combined across
        implicates.

    Notes
    -----
    Base R's lm()/glm() only provide classical (non-robust) standard errors
    - there's no HC0-3 equivalent without the 'sandwich' package, which this
    adapter deliberately doesn't pull in (keeping the R-side dependency at
    just R itself). Use `r_fixest_adapter` for built-in robust/cluster SEs.
    """
    from . import _r_interop as _r

    fn = "glm" if family else "lm"
    fit_code = _r.call(
        fn,
        _r.RRaw(formula),
        data=_r.RRaw("{df}"),
        weights=_r.RRaw(weight) if weight else None,
        family=family,
        **r_kwargs,
    )

    coef, vcov, tidy = _r.fit_r_model(
        df, fit_code, tidy_code="summary({fit})$coefficients"
    )

    df_estimates = _r.coef_table(coef, join_on_name, value_name)
    df_ses = _r.ses_from_vcov(vcov, join_on_name, value_name)
    df_vcov = _r.vcov_table(vcov, join_on_name, value_name)
    df_tidy = _r.matrix_table(tidy, join_on_name)

    return _adapter_stats(df_estimates, df_ses, df_vcov, df_tidy, join_on_name)

r_fixest_adapter

r_fixest_adapter(
    df,
    formula: str,
    func: str = "feols",
    weight: str | None = None,
    family: str | None = None,
    vcov: str | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats

Fit any fixest regression and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape. fixest supports fixed effects directly in the formula (e.g. "y ~ x1 | firm + year") and computes robust/clustered SEs natively - no 'sandwich' needed. Requires the R 'fixest' package - see survey_kit.statistics._r_interop.check_r_setup(["fixest"]).

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
formula fixest formula string, e.g. "y ~ x1 + x2 | firm" for a fit

with a firm fixed effect.

required
func which fixest estimator to call, e.g. "feols" (default), "feglm",

"fepois", "femlm", "feNmlm", "feglm.fit" - anything in the fixest namespace. Called as fixest::{func}(...).

'feols'
weight column name for weighted estimation, or None. Default is None.

Passed as fixest's weights=~column (a one-sided formula - unlike most other arguments here, this can't be a plain quoted string).

None
family a plain R family name (e.g. "binomial", "poisson") - only

meaningful for func="feglm"/"femlm". For a non-default link function, wrap the full expression in _r_interop.RRaw. Default is None.

None
vcov fixest's own `vcov=` argument - a string like "hetero" (robust)

or "iid" (classical), or a one-sided formula string like "~firm" for cluster-robust SEs. Defaults to "hetero" for the same reason statsmodels_adapter defaults to HC3 - rarely safe to assume homoskedasticity - unless you pass cluster= via r_kwargs, in which case vcov is left unset so fixest can infer clustered SEs from it: fixest raises if vcov and cluster are both given.

None
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE/covariance value column in the

output. Default is "estimate".

'estimate'
**r_kwargs any other argument any fixest estimator takes - `cluster`,

panel.id, split/fsplit, se, ssc, lean, notes, verbose, etc. - converted to R via _r_interop.py_to_r_literal: None omits the argument, bool/int/float/str/list/dict convert naturally (a plain column name like cluster="firm" works exactly as it does when you type it directly in fixest), a string starting with "~" is passed through raw (a formula, e.g. panel.id="~id+time"), and _r_interop.RRaw wraps anything else that needs to be emitted as literal R code (e.g. ssc=RRaw('ssc(fixef.K="full")')). R argument names with a "." (e.g. panel.id) can't be Python keyword names directly - build the kwargs dict separately and splat it: **{"panel.id": "~id+time"}.

{}

Returns:

Type Description
AdapterStats

(df_estimates, df_ses, df_vcov) - df_vcov is always populated here since vcov() is free alongside coef() in fixest's results.

Source code in src/survey_kit/statistics/adapters.py
def r_fixest_adapter(
    df,
    formula: str,
    func: str = "feols",
    weight: str | None = None,
    family: str | None = None,
    vcov: str | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats:
    """
    Fit any fixest regression and return its coefficient table in
    survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.
    fixest supports fixed effects directly in the formula (e.g. "y ~ x1 |
    firm + year") and computes robust/clustered SEs natively - no 'sandwich'
    needed. Requires the R 'fixest' package - see
    `survey_kit.statistics._r_interop.check_r_setup(["fixest"])`.

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    formula : fixest formula string, e.g. "y ~ x1 + x2 | firm" for a fit
        with a firm fixed effect.
    func : which fixest estimator to call, e.g. "feols" (default), "feglm",
        "fepois", "femlm", "feNmlm", "feglm.fit" - anything in the fixest
        namespace. Called as `fixest::{func}(...)`.
    weight : column name for weighted estimation, or None. Default is None.
        Passed as fixest's `weights=~column` (a one-sided formula - unlike
        most other arguments here, this can't be a plain quoted string).
    family : a plain R family name (e.g. "binomial", "poisson") - only
        meaningful for func="feglm"/"femlm". For a non-default link
        function, wrap the full expression in `_r_interop.RRaw`. Default
        is None.
    vcov : fixest's own `vcov=` argument - a string like "hetero" (robust)
        or "iid" (classical), or a one-sided formula string like "~firm"
        for cluster-robust SEs. Defaults to "hetero" for the same reason
        statsmodels_adapter defaults to HC3 - rarely safe to assume
        homoskedasticity - *unless* you pass `cluster=` via r_kwargs, in
        which case vcov is left unset so fixest can infer clustered SEs
        from it: fixest raises if vcov and cluster are both given.
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE/covariance value column in the
        output. Default is "estimate".
    **r_kwargs : any other argument any fixest estimator takes - `cluster`,
        `panel.id`, `split`/`fsplit`, `se`, `ssc`, `lean`, `notes`,
        `verbose`, etc. - converted to R via `_r_interop.py_to_r_literal`:
        None omits the argument, bool/int/float/str/list/dict convert
        naturally (a plain column name like `cluster="firm"` works exactly
        as it does when you type it directly in fixest), a string starting
        with "~" is passed through raw (a formula, e.g.
        `panel.id="~id+time"`), and `_r_interop.RRaw` wraps anything else
        that needs to be emitted as literal R code (e.g.
        `ssc=RRaw('ssc(fixef.K="full")')`). R argument names with a "."
        (e.g. `panel.id`) can't be Python keyword names directly - build the
        kwargs dict separately and splat it: `**{"panel.id": "~id+time"}`.

    Returns
    -------
    AdapterStats
        (df_estimates, df_ses, df_vcov) - df_vcov is always populated here
        since vcov() is free alongside coef() in fixest's results.
    """
    return _fixest_fit(
        df, func, formula, weight, family, vcov, join_on_name, value_name, r_kwargs
    )

r_feols

r_feols(
    df,
    formula: str,
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats
fixest::feols() (linear regression, with fixed effects) with arguments
mirroring fixest's own. See `r_fixest_adapter` for the fully generic
`func=` version (feNmlm, feglm.fit, ...) and the general conversion
rules; this is the same thing with feols's common arguments spelled out.
Parameters
df : the merged implicate data (supplied by mi_ses_from_function).
formula : fixest formula string, e.g. "y ~ x1 + x2 | firm" for a fit
    with a firm fixed effect.
weight : column name for weighted estimation, or None. Passed as
    `weights=~column` (fixest needs a formula here, not a plain string).
vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
    formula string like "~firm" for cluster-robust SEs.

{params} join_on_name : name of the term-identifier column in the output. Default is "Variable". value_name : name of the coefficient/SE/covariance value column in the output. Default is "estimate".

Returns
AdapterStats
    (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
    coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
    snapshot only, never combined across implicates.
Source code in src/survey_kit/statistics/adapters.py
def r_feols(
    df,
    formula: str,
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats:
    """
        fixest::feols() (linear regression, with fixed effects) with arguments
        mirroring fixest's own. See `r_fixest_adapter` for the fully generic
        `func=` version (feNmlm, feglm.fit, ...) and the general conversion
        rules; this is the same thing with feols's common arguments spelled out.

        Parameters
        ----------
        df : the merged implicate data (supplied by mi_ses_from_function).
        formula : fixest formula string, e.g. "y ~ x1 + x2 | firm" for a fit
            with a firm fixed effect.
        weight : column name for weighted estimation, or None. Passed as
            `weights=~column` (fixest needs a formula here, not a plain string).
        vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
            formula string like "~firm" for cluster-robust SEs.
    {params}
        join_on_name : name of the term-identifier column in the output.
            Default is "Variable".
        value_name : name of the coefficient/SE/covariance value column in the
            output. Default is "estimate".

        Returns
        -------
        AdapterStats
            (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
            coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
            snapshot only, never combined across implicates.
    """
    r_kwargs = _fixest_named_kwargs(
        cluster, panel_id, ssc, fixef, lean, notes, verbose, r_kwargs
    )
    return _fixest_fit(
        df, "feols", formula, weight, None, vcov, join_on_name, value_name, r_kwargs
    )

r_feglm

r_feglm(
    df,
    formula: str,
    family: str = "gaussian",
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats
fixest::feglm() (GLM, with fixed effects) with arguments mirroring
fixest's own. See `r_fixest_adapter` for the fully generic `func=`
version and the general conversion rules.
Parameters
df : the merged implicate data (supplied by mi_ses_from_function).
formula : fixest formula string, e.g. "y ~ x1 + x2 | firm".
family : a plain R family name, e.g. "binomial", "poisson" - passed as a
    quoted string (fixest resolves it to the family function itself).
    For a non-default link function, wrap the full expression in
    [`RRaw`][survey_kit.statistics._r_interop.RRaw], e.g.
    `family=RRaw('binomial(link="probit")')`. Default is "gaussian"
    (matching feglm's own default) - for a pure Poisson fit,
    `r_fepois` is faster and doesn't need this.
weight : column name for weighted estimation, or None. Passed as
    `weights=~column`.
vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
    formula string like "~firm" for cluster-robust SEs.

{params} join_on_name : name of the term-identifier column in the output. Default is "Variable". value_name : name of the coefficient/SE/covariance value column in the output. Default is "estimate".

Returns
AdapterStats
    (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
    coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
    snapshot only, never combined across implicates.
Source code in src/survey_kit/statistics/adapters.py
def r_feglm(
    df,
    formula: str,
    family: str = "gaussian",
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats:
    """
        fixest::feglm() (GLM, with fixed effects) with arguments mirroring
        fixest's own. See `r_fixest_adapter` for the fully generic `func=`
        version and the general conversion rules.

        Parameters
        ----------
        df : the merged implicate data (supplied by mi_ses_from_function).
        formula : fixest formula string, e.g. "y ~ x1 + x2 | firm".
        family : a plain R family name, e.g. "binomial", "poisson" - passed as a
            quoted string (fixest resolves it to the family function itself).
            For a non-default link function, wrap the full expression in
            [`RRaw`][survey_kit.statistics._r_interop.RRaw], e.g.
            `family=RRaw('binomial(link="probit")')`. Default is "gaussian"
            (matching feglm's own default) - for a pure Poisson fit,
            `r_fepois` is faster and doesn't need this.
        weight : column name for weighted estimation, or None. Passed as
            `weights=~column`.
        vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
            formula string like "~firm" for cluster-robust SEs.
    {params}
        join_on_name : name of the term-identifier column in the output.
            Default is "Variable".
        value_name : name of the coefficient/SE/covariance value column in the
            output. Default is "estimate".

        Returns
        -------
        AdapterStats
            (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
            coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
            snapshot only, never combined across implicates.
    """
    r_kwargs = _fixest_named_kwargs(
        cluster, panel_id, ssc, fixef, lean, notes, verbose, r_kwargs
    )
    return _fixest_fit(
        df, "feglm", formula, weight, family, vcov, join_on_name, value_name, r_kwargs
    )

r_fepois

r_fepois(
    df,
    formula: str,
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats
fixest::fepois() (Poisson regression, with fixed effects) with arguments
mirroring fixest's own. See `r_fixest_adapter` for the fully generic
`func=` version and the general conversion rules.
Parameters
df : the merged implicate data (supplied by mi_ses_from_function).
formula : fixest formula string, e.g. "y ~ x1 + x2 | firm".
weight : column name for weighted estimation, or None. Passed as
    `weights=~column`.
vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
    formula string like "~firm" for cluster-robust SEs.

{params} join_on_name : name of the term-identifier column in the output. Default is "Variable". value_name : name of the coefficient/SE/covariance value column in the output. Default is "estimate".

Returns
AdapterStats
    (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
    coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
    snapshot only, never combined across implicates.
Source code in src/survey_kit/statistics/adapters.py
def r_fepois(
    df,
    formula: str,
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats:
    """
        fixest::fepois() (Poisson regression, with fixed effects) with arguments
        mirroring fixest's own. See `r_fixest_adapter` for the fully generic
        `func=` version and the general conversion rules.

        Parameters
        ----------
        df : the merged implicate data (supplied by mi_ses_from_function).
        formula : fixest formula string, e.g. "y ~ x1 + x2 | firm".
        weight : column name for weighted estimation, or None. Passed as
            `weights=~column`.
        vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
            formula string like "~firm" for cluster-robust SEs.
    {params}
        join_on_name : name of the term-identifier column in the output.
            Default is "Variable".
        value_name : name of the coefficient/SE/covariance value column in the
            output. Default is "estimate".

        Returns
        -------
        AdapterStats
            (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
            coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
            snapshot only, never combined across implicates.
    """
    r_kwargs = _fixest_named_kwargs(
        cluster, panel_id, ssc, fixef, lean, notes, verbose, r_kwargs
    )
    return _fixest_fit(
        df, "fepois", formula, weight, None, vcov, join_on_name, value_name, r_kwargs
    )

r_femlm

r_femlm(
    df,
    formula: str,
    family: str = "poisson",
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats
fixest::femlm() (max-likelihood: Poisson/negative binomial/logit/
Gaussian, with fixed effects) with arguments mirroring fixest's own.
See `r_fixest_adapter` for the fully generic `func=` version and the
general conversion rules. Note: femlm has no `weights=` argument
(unlike feols/feglm/fepois).
Parameters
df : the merged implicate data (supplied by mi_ses_from_function).
formula : fixest formula string, e.g. "y ~ x1 + x2 | firm".
family : one of "poisson" (default), "negbin", "logit", "gaussian" -
    a plain string (femlm, unlike feglm, only accepts one of these four
    exact names - there's no link-function customization here).
vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
    formula string like "~firm" for cluster-robust SEs.

{params} join_on_name : name of the term-identifier column in the output. Default is "Variable". value_name : name of the coefficient/SE/covariance value column in the output. Default is "estimate".

Returns
AdapterStats
    (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
    coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
    snapshot only, never combined across implicates.
Source code in src/survey_kit/statistics/adapters.py
def r_femlm(
    df,
    formula: str,
    family: str = "poisson",
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    **r_kwargs,
) -> AdapterStats:
    """
        fixest::femlm() (max-likelihood: Poisson/negative binomial/logit/
        Gaussian, with fixed effects) with arguments mirroring fixest's own.
        See `r_fixest_adapter` for the fully generic `func=` version and the
        general conversion rules. Note: femlm has no `weights=` argument
        (unlike feols/feglm/fepois).

        Parameters
        ----------
        df : the merged implicate data (supplied by mi_ses_from_function).
        formula : fixest formula string, e.g. "y ~ x1 + x2 | firm".
        family : one of "poisson" (default), "negbin", "logit", "gaussian" -
            a plain string (femlm, unlike feglm, only accepts one of these four
            exact names - there's no link-function customization here).
        vcov : "hetero" (robust, the default), "iid" (classical), or a one-sided
            formula string like "~firm" for cluster-robust SEs.
    {params}
        join_on_name : name of the term-identifier column in the output.
            Default is "Variable".
        value_name : name of the coefficient/SE/covariance value column in the
            output. Default is "estimate".

        Returns
        -------
        AdapterStats
            (df_estimates, df_ses, df_vcov, df_tidy) - df_tidy is fixest's own
            coeftable() (Estimate/Std. Error/t value/Pr(>|t|)), a diagnostic
            snapshot only, never combined across implicates.
    """
    r_kwargs = _fixest_named_kwargs(
        cluster, panel_id, ssc, fixef, lean, notes, verbose, r_kwargs
    )
    return _fixest_fit(
        df, "femlm", formula, None, family, vcov, join_on_name, value_name, r_kwargs
    )

mi_ses_from_r_fixest

Namespace for per-estimator mi_ses_from_function(delegate=r_feols/ r_feglm/r_fepois/r_femlm, ...) wrappers - mi_ses_from_r_fixest.feols(...), .feglm(...), .fepois(...), .femlm(...) - each naming that estimator's own arguments (formula, weight, vcov, cluster, panel_id, ...) directly instead of packing them into an arguments={} dict, so IDEs show the right parameters for the one you're actually calling. Everything else the R side of fixest accepts still flows through **r_kwargs exactly as in r_feols/r_feglm/ r_fepois/r_femlm. See r_fixest_adapter for the fully generic func= escape hatch (feNmlm, feglm.fit, ...) this doesn't cover.

feglm staticmethod

feglm(
    df_implicates,
    formula: str,
    family: str = "gaussian",
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **r_kwargs,
)

mi_ses_from_function(delegate=r_feglm, ...) - fixest::feglm() (GLM, with fixed effects) across implicates.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
formula str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
family str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
weight str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
vcov str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
cluster str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
panel_id str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
ssc str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
fixef str

lean, notes, verbose, join_on_name, value_name, **r_kwargs : see r_feglm.

required
replicates see `mi_ses_from_r_fixest.feols` - same

replicate-weight-bootstrap option, vcov forced to "iid", cluster forced off, and weight ignored in this mode. Default is None.

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

r_feglm : the delegate this wraps.

femlm staticmethod

femlm(
    df_implicates,
    formula: str,
    family: str = "poisson",
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **r_kwargs,
)

mi_ses_from_function(delegate=r_femlm, ...) - fixest::femlm() (max-likelihood: Poisson/negative binomial/logit/Gaussian, with fixed effects) across implicates. Note: femlm has no weight= argument (unlike feols/feglm/fepois), so - unlike those three - there's no replicates= option here either: nothing to substitute a replicate weight column into.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
formula str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
family str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
vcov str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
cluster str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
panel_id str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
ssc str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
fixef str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
lean str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required
notes str

verbose, join_on_name, value_name, **r_kwargs : see r_femlm.

required

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

r_femlm : the delegate this wraps.

feols staticmethod

feols(
    df_implicates,
    formula: str,
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **r_kwargs,
)

mi_ses_from_function(delegate=r_feols, ...) - fixest::feols() (linear regression, with fixed effects) across implicates.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
formula str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
weight str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
vcov str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
cluster str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
panel_id str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
ssc str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
fixef str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
lean str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
notes str

verbose, join_on_name, value_name, **r_kwargs : see r_feols.

required
replicates `survey_kit.statistics.replicates.Replicates`,

optional. When given, per-implicate SEs come from resampling across replicate weights instead of vcov/cluster - r_feols is run once per replicate weight column (point estimates only, with vcov forced to "iid" and cluster forced off since they're discarded anyway), the same replicate-weight-bootstrap approach mi_ses_from_stata's replicates= uses. weight is ignored in this mode; the implicate is converted to an R data.frame once, not once per replicate. Default is None (use r_feols's own vcov/cluster).

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

r_feols : the delegate this wraps.

fepois staticmethod

fepois(
    df_implicates,
    formula: str,
    weight: str | None = None,
    vcov: str | None = None,
    cluster=None,
    panel_id=None,
    ssc=None,
    fixef=None,
    lean: bool | None = None,
    notes: bool | None = None,
    verbose: int | None = None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    replicates=None,
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
    **r_kwargs,
)

mi_ses_from_function(delegate=r_fepois, ...) - fixest::fepois() (Poisson regression, with fixed effects) across implicates.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
formula str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
weight str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
vcov str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
cluster str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
panel_id str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
ssc str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
fixef str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
lean str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
notes str

verbose, join_on_name, value_name, **r_kwargs : see r_fepois.

required
replicates see `mi_ses_from_r_fixest.feols` - same

replicate-weight-bootstrap option, vcov forced to "iid", cluster forced off, and weight ignored in this mode. Default is None.

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

See Also

r_fepois : the delegate this wraps.

Stata (pystata)

Requires Stata 17+ plus pip install survey-kit[stata]. stata_adapter (and its mi_ses_from_stata shortcut) runs any e-class command as a plain string - there's no separate named-wrapper-per-estimator layer to route around the way fixest has, since any Stata command already works by just changing the command string. stata_results_adapter reaches any r()/e() result rather than the fixed e(b)/e(V)/r(table) triplet, and is what mi_ses_from_stata's replicates= option uses under the hood for replicate-weight bootstrapping.

stata_adapter

stata_adapter(
    df,
    command: str | list[str],
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    edition: str | None = None,
    stata_path: str | None = None,
    reuse_data: bool = False,
    quietly: bool = True,
) -> AdapterStats

Run an arbitrary Stata e-class estimation command (regress, logit, xtreg, areg, svy: ..., or anything from an installed community package) and return its coefficient table in survey_kit's normalized (df_estimates, df_ses, df_vcov, df_tidy) shape.

See survey_kit.statistics._stata_interop's module docstring for details, and check_stata_setup() there for a setup diagnostic.

Data moves into Stata via a .dta file written by polars_readstat (write_readstat) rather than pystata's own DataFrame transfer - .dta is Stata's own native, most battle-tested ingestion path, so this doesn't depend on however pystata's transfer mechanism behaves.

Parameters:

Name Type Description Default
df the merged implicate data (supplied by mi_ses_from_function).
required
command the Stata command to run, e.g. "regress y x1 x2", or

"regress y x1 x2 [pw=w]", or "xtreg y x1 x2, fe". Must leave e(b)/e(V) populated - true of most estimation commands. Pass a list of commands run in order, instead of a single string, when you need setup (gen/svyset/...) between use and the estimation command - only the LAST command's e(b)/e(V)/r(table) are read back, e.g. ["svyset psu [pw=weight], strata(strata)", "svy: regress y x1 x2"].

required
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the coefficient/SE/covariance value column in the

output. Default is "estimate".

'estimate'
edition forwarded to

_stata_interop.require_pystata - stata_path (your Stata install directory) is only needed if pystata's utilities folder isn't already on sys.path; edition ("be"/"se"/"mp") only if it can't be auto-detected. Only used on the first call in a process - Stata stays initialized afterward, like a package import.

None
stata_path forwarded to

_stata_interop.require_pystata - stata_path (your Stata install directory) is only needed if pystata's utilities folder isn't already on sys.path; edition ("be"/"se"/"mp") only if it can't be auto-detected. Only used on the first call in a process - Stata stays initialized afterward, like a package import.

None
reuse_data if True, skip re-exporting/re-`use`-ing df when it's the

same object (by identity) as a previous reuse_data=True call - useful when calling this repeatedly for the same underlying data (e.g. once per replicate weight, if command already references the weight column directly rather than needing a different df per call). Default is False. Call survey_kit.statistics._stata_interop.clear_stata_cache() once done with data reused this way, to free Stata's own copy.

False
quietly pass False to let `command` stream Stata's own console

output on success too (failures always surface the real error text regardless - see _stata_interop._run_in_stata's docstring). Default is True.

True

Returns:

Type Description
AdapterStats

(df_estimates, df_ses, df_vcov, df_tidy) - df_vcov from e(V). df_tidy is Stata's own r(table) (the matrix ci/test/etc. use internally: b/se/t-or-z/p-value/CI, transposed to one row per term) populated automatically after any estimation command - a diagnostic snapshot only, never combined across implicates.

Source code in src/survey_kit/statistics/adapters.py
def stata_adapter(
    df,
    command: str | list[str],
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    edition: str | None = None,
    stata_path: str | None = None,
    reuse_data: bool = False,
    quietly: bool = True,
) -> AdapterStats:
    """
    Run an arbitrary Stata e-class estimation command (regress, logit,
    xtreg, areg, svy: ..., or anything from an installed community package)
    and return its coefficient table in survey_kit's normalized
    (df_estimates, df_ses, df_vcov, df_tidy) shape.

    See `survey_kit.statistics._stata_interop`'s module docstring for
    details, and `check_stata_setup()` there for a setup diagnostic.

    Data moves into Stata via a .dta file written by polars_readstat
    (`write_readstat`) rather than pystata's own DataFrame transfer - `.dta`
    is Stata's own native, most battle-tested ingestion path, so this
    doesn't depend on however pystata's transfer mechanism behaves.

    Parameters
    ----------
    df : the merged implicate data (supplied by mi_ses_from_function).
    command : the Stata command to run, e.g. "regress y x1 x2", or
        "regress y x1 x2 [pw=w]", or "xtreg y x1 x2, fe". Must leave
        e(b)/e(V) populated - true of most estimation commands. Pass a
        list of commands run in order, instead of a single string, when
        you need setup (`gen`/`svyset`/...) between `use` and the
        estimation command - only the LAST command's e(b)/e(V)/r(table)
        are read back, e.g.
        `["svyset psu [pw=weight], strata(strata)", "svy: regress y x1 x2"]`.
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the coefficient/SE/covariance value column in the
        output. Default is "estimate".
    edition, stata_path : forwarded to
        `_stata_interop.require_pystata` - stata_path (your Stata install
        directory) is only needed if pystata's utilities folder isn't
        already on sys.path; edition ("be"/"se"/"mp") only if it can't be
        auto-detected. Only used on the first call in a
        process - Stata stays initialized afterward, like a package import.
    reuse_data : if True, skip re-exporting/re-`use`-ing df when it's the
        same object (by identity) as a previous reuse_data=True call -
        useful when calling this repeatedly for the same underlying data
        (e.g. once per replicate weight, if `command` already references
        the weight column directly rather than needing a different df per
        call). Default is False. Call
        `survey_kit.statistics._stata_interop.clear_stata_cache()` once
        done with data reused this way, to free Stata's own copy.
    quietly : pass False to let `command` stream Stata's own console
        output on success too (failures always surface the real error
        text regardless - see `_stata_interop._run_in_stata`'s
        docstring). Default is True.

    Returns
    -------
    AdapterStats
        (df_estimates, df_ses, df_vcov, df_tidy) - df_vcov from e(V).
        df_tidy is Stata's own r(table) (the matrix `ci`/`test`/etc. use
        internally: b/se/t-or-z/p-value/CI, transposed to one row per term)
        populated automatically after any estimation command - a
        diagnostic snapshot only, never combined across implicates.
    """
    from . import _stata_interop as _st

    b, b_names, V, table, table_row_names, table_col_names = _st.run_stata_model(
        df,
        command,
        edition=edition,
        stata_path=stata_path,
        reuse_data=reuse_data,
        quietly=quietly,
    )

    df_estimates = pl.DataFrame({join_on_name: b_names, value_name: b})

    import numpy as np

    v_arr = np.asarray(V)
    se = np.sqrt(np.diag(v_arr)).tolist()
    df_ses = pl.DataFrame({join_on_name: b_names, value_name: se})

    n = len(b_names)
    df_vcov = pl.DataFrame(
        [
            {
                f"{join_on_name}_1": b_names[i],
                f"{join_on_name}_2": b_names[j],
                value_name: float(v_arr[i, j]),
            }
            for i in range(n)
            for j in range(n)
        ]
    )

    #   r(table) is stat-by-term (rows=stats, cols=terms) - transpose to
    #   the one-row-per-term shape every other adapter's df_tidy uses.
    table_arr = np.asarray(table)
    tidy_data = {join_on_name: table_col_names}
    for i, stat_name in enumerate(table_row_names):
        tidy_data[stat_name] = table_arr[i, :].tolist()
    df_tidy = pl.DataFrame(tidy_data)

    return _adapter_stats(df_estimates, df_ses, df_vcov, df_tidy, join_on_name)

mi_ses_from_stata

mi_ses_from_stata(
    df_implicates,
    command: str | list[str],
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    edition: str | None = None,
    stata_path: str | None = None,
    reuse_data: bool = False,
    quietly: bool = True,
    replicates=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
)

mi_ses_from_function(delegate=stata_adapter, ...), with stata_adapter's own arguments (command, edition, stata_path, ...) taken directly as keyword arguments instead of packed into an arguments={} dict - the common case of running one Stata command across implicates.

Parameters:

Name Type Description Default
df_implicates

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
path_srmi

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
index

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
df_noimputes

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
parallel

parallel_inputs, rounding, round_output : see mi_ses_from_function.

required
command str | list[str]

quietly : see stata_adapter. When replicates is given, command needs a "{weight}" placeholder instead (see below) - reuse_data is then forced on internally regardless of what's passed here (see replicates).

required
join_on_name str | list[str]

quietly : see stata_adapter. When replicates is given, command needs a "{weight}" placeholder instead (see below) - reuse_data is then forced on internally regardless of what's passed here (see replicates).

required
value_name str | list[str]

quietly : see stata_adapter. When replicates is given, command needs a "{weight}" placeholder instead (see below) - reuse_data is then forced on internally regardless of what's passed here (see replicates).

required
edition str | list[str]

quietly : see stata_adapter. When replicates is given, command needs a "{weight}" placeholder instead (see below) - reuse_data is then forced on internally regardless of what's passed here (see replicates).

required
stata_path str | list[str]

quietly : see stata_adapter. When replicates is given, command needs a "{weight}" placeholder instead (see below) - reuse_data is then forced on internally regardless of what's passed here (see replicates).

required
reuse_data str | list[str]

quietly : see stata_adapter. When replicates is given, command needs a "{weight}" placeholder instead (see below) - reuse_data is then forced on internally regardless of what's passed here (see replicates).

required
replicates `survey_kit.statistics.replicates.Replicates`, optional.

When given, per-implicate SEs come from resampling across replicate weights instead of command's own e(V) - command is run once per replicate weight column (via stata_results_adapter, reading back e(b) only) rather than Stata's own bootstrap/brr/ jackknife prefix, so it needs a "{weight}" placeholder, e.g. "regress y x1 x2 [pw={weight}]". Stata's own replicate-estimation commands are usually faster and more idiomatic if you're already set up for them - this is for reusing survey_kit's own Replicates/StatCalculator machinery (e.g. to match SEs computed the same way elsewhere in a project) instead. Default is None (use command's own e(V), the stata_adapter path).

None

Returns:

Type Description
MultipleImputation

Same as mi_ses_from_function.

Examples:

>>> mi_reg = mi_ses_from_stata(
...     df_implicates=df_implicates,
...     command="regress y x1 x2",
... )

Replicate-weight SEs instead of Stata's own e(V):

>>> from survey_kit.statistics.replicates import Replicates
>>> mi_reg = mi_ses_from_stata(
...     df_implicates=df_implicates,
...     command="regress y x1 x2 [pw={weight}]",
...     replicates=Replicates(weight_stub="replicate_", n_replicates=80),
... )
See Also

stata_adapter, stata_results_adapter : the delegates this wraps. mi_ses_from_function : the general-purpose function this specializes.

Source code in src/survey_kit/statistics/adapters.py
def mi_ses_from_stata(
    df_implicates,
    command: str | list[str],
    path_srmi: str = "",
    index: list | None = None,
    df_noimputes=None,
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    edition: str | None = None,
    stata_path: str | None = None,
    reuse_data: bool = False,
    quietly: bool = True,
    replicates=None,
    parallel: bool = False,
    parallel_inputs=None,
    rounding=None,
    round_output: bool = True,
):
    """
    `mi_ses_from_function(delegate=stata_adapter, ...)`, with
    `stata_adapter`'s own arguments (`command`, `edition`, `stata_path`,
    ...) taken directly as keyword arguments instead of packed into an
    `arguments={}` dict - the common case of running one Stata command
    across implicates.

    Parameters
    ----------
    df_implicates, path_srmi, index, df_noimputes, parallel,
        parallel_inputs, rounding, round_output : see
        [`mi_ses_from_function`][survey_kit.statistics.multiple_imputation.mi_ses_from_function].
    command, join_on_name, value_name, edition, stata_path, reuse_data,
        quietly : see `stata_adapter`. When `replicates` is given, `command`
        needs a "{weight}" placeholder instead (see below) - `reuse_data`
        is then forced on internally regardless of what's passed here (see
        `replicates`).
    replicates : `survey_kit.statistics.replicates.Replicates`, optional.
        When given, per-implicate SEs come from resampling across
        replicate weights instead of `command`'s own e(V) - `command` is
        run once per replicate weight column (via `stata_results_adapter`,
        reading back e(b) only) rather than Stata's own `bootstrap`/`brr`/
        `jackknife` prefix, so it needs a "{weight}" placeholder, e.g.
        `"regress y x1 x2 [pw={weight}]"`. Stata's own replicate-estimation
        commands are usually faster and more idiomatic if you're already
        set up for them - this is for reusing survey_kit's own
        Replicates/StatCalculator machinery (e.g. to match SEs computed
        the same way elsewhere in a project) instead. Default is None (use
        `command`'s own e(V), the `stata_adapter` path).

    Returns
    -------
    MultipleImputation
        Same as `mi_ses_from_function`.

    Examples
    --------
    >>> mi_reg = mi_ses_from_stata(
    ...     df_implicates=df_implicates,
    ...     command="regress y x1 x2",
    ... )

    Replicate-weight SEs instead of Stata's own e(V):

    >>> from survey_kit.statistics.replicates import Replicates
    >>> mi_reg = mi_ses_from_stata(
    ...     df_implicates=df_implicates,
    ...     command="regress y x1 x2 [pw={weight}]",
    ...     replicates=Replicates(weight_stub="replicate_", n_replicates=80),
    ... )

    See Also
    --------
    stata_adapter, stata_results_adapter : the delegates this wraps.
    [`mi_ses_from_function`][survey_kit.statistics.multiple_imputation.mi_ses_from_function] :
        the general-purpose function this specializes.
    """
    from .calculator import StatCalculator
    from .multiple_imputation import mi_ses_from_function
    from . import _stata_interop as _st

    if replicates is None:
        delegate = stata_adapter
        arguments = {
            "command": command,
            "join_on_name": join_on_name,
            "value_name": value_name,
            "edition": edition,
            "stata_path": stata_path,
            "reuse_data": reuse_data,
            "quietly": quietly,
        }
    else:

        def delegate(df):
            calc = StatCalculator.from_function(
                delegate=stata_results_adapter,
                estimate_ids=[join_on_name],
                df=df,
                arguments={
                    "command": command,
                    "results": ["e(b)"],
                    "join_on_name": join_on_name,
                    "value_name": value_name,
                    "edition": edition,
                    "stata_path": stata_path,
                    #   the same implicate df is reused across every
                    #   replicate weight in this loop - only `weight`
                    #   changes - so re-exporting/re-`use`-ing it every
                    #   call would be pure waste.
                    "reuse_data": True,
                    "quietly": quietly,
                },
                replicates=replicates,
                display=False,
            )
            _st.clear_stata_cache()
            return calc

        arguments = {}

    return mi_ses_from_function(
        delegate=delegate,
        df_implicates=df_implicates,
        join_on=[join_on_name],
        path_srmi=path_srmi,
        index=index,
        df_noimputes=df_noimputes,
        arguments=arguments,
        parallel=parallel,
        parallel_inputs=parallel_inputs,
        rounding=rounding,
        round_output=round_output,
    )

stata_results_adapter

stata_results_adapter(
    df,
    command: str | list[str],
    results: list[str],
    weight: str = "",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    edition: str | None = None,
    stata_path: str | None = None,
    reuse_data: bool = False,
    quietly: bool = True,
) -> pl.DataFrame

Run an arbitrary Stata command (r-class or e-class) and return a flat table of exactly the r()/e() results you name - the Stata counterpart to survey_kit's generic R/rpy2 escape hatch (get_library/ extract_fit in _r_interop.py), but shaped as a StatCalculator.from_function delegate (df, weight -> one row per estimate) rather than an mi_ses_from_function delegate: point estimates only, no vcov. Use this for bootstrap/replicate-weight variance - the spread of estimates across replicate weights IS the SE (computed by Replicates/StatCalculator), not Stata's own e(V) - the same reason the R tutorial's ad-hoc run_regression delegate returns just a plain coefficient table with no SE of its own.

Unlike stata_adapter (which assumes an e-class fit and always reads the fixed e(b)/e(V)/r(table) triplet), this works for ANY command that populates r()/e() results - summarize, tabstat, ci, svy: mean, or a full e-class regression - you just name what you want back.

Parameters:

Name Type Description Default
df one replicate/bootstrap draw's data (supplied by

StatCalculator.from_function per replicate weight).

required
command the Stata command to run, with "{weight}" as a placeholder

for the weight column name if weight is given, e.g. "regress y x1 x2 [pw={weight}]" or "summarize x [aw={weight}]". Stata's weight syntax/placement varies enough by command (pw/aw/ fw/iw, and where the bracket goes) that this doesn't try to build it for you - write it the way you'd type it directly in Stata. Pass a list of commands run in order, instead of a single string, when you need setup (a svyset, say) between use and the command whose results you actually want back - each element may use "{weight}", e.g. ["svyset psu [pw={weight}]", "svy: mean x"].

required
results names to pull back, e.g. ["r(mean)", "r(Var)"] or ["e(b)"].

A scalar becomes one row (join_on_name=name); a matrix's columns (e.g. e(b)'s term names) each become their own row.

required
weight column name to substitute into "{weight}" in `command`, or ""

(the StatCalculator.from_function convention - see its weight_argument_name) if command needs no weight or already hardcodes one. Default is "".

''
join_on_name name of the term-identifier column in the output.

Default is "Variable".

'Variable'
value_name name of the estimate column in the output. Default is

"estimate".

'estimate'
edition forwarded to

_stata_interop.require_pystata.

None
stata_path forwarded to

_stata_interop.require_pystata.

None
reuse_data if True, skip re-exporting/re-`use`-ing df when it's the

same object (by identity) as the previous call - this is exactly the common case here: StatCalculator.from_function calls this delegate once per replicate weight with the same df object every time (only weight, spliced into command, changes), so the underlying data doesn't need re-exporting on every replicate. Default is False (safe/explicit opt-in - see _stata_interop._run_in_stata's docstring for why this can't be a default the way the R side's dataframe_to_r caching is). Call survey_kit.statistics._stata_interop.clear_stata_cache() once the replicate loop is done, to free Stata's own copy of the data.

False
quietly pass False to let `command` stream Stata's own console

output on success too (failures always surface the real error text regardless - see _stata_interop._run_in_stata's docstring). Default is True.

True

Returns:

Type Description
DataFrame

One row per named scalar, or per column of a named matrix - exactly the shape a StatCalculator.from_function delegate needs (no vcov/tidy - see the note above on where the variance actually comes from).

Source code in src/survey_kit/statistics/adapters.py
def stata_results_adapter(
    df,
    command: str | list[str],
    results: list[str],
    weight: str = "",
    join_on_name: str = "Variable",
    value_name: str = "estimate",
    edition: str | None = None,
    stata_path: str | None = None,
    reuse_data: bool = False,
    quietly: bool = True,
) -> pl.DataFrame:
    """
    Run an arbitrary Stata command (r-class or e-class) and return a flat
    table of exactly the r()/e() results you name - the Stata counterpart
    to survey_kit's generic R/rpy2 escape hatch (`get_library`/
    `extract_fit` in `_r_interop.py`), but shaped as a
    [`StatCalculator.from_function`][survey_kit.statistics.calculator.StatCalculator.from_function]
    delegate (df, weight -> one row per estimate) rather than an
    `mi_ses_from_function` delegate: point estimates only, no vcov. Use
    this for bootstrap/replicate-weight variance - the spread of estimates
    across replicate weights IS the SE (computed by
    `Replicates`/`StatCalculator`), not Stata's own e(V) - the same reason
    the R tutorial's ad-hoc `run_regression` delegate returns just a plain
    coefficient table with no SE of its own.

    Unlike `stata_adapter` (which assumes an e-class fit and always reads
    the fixed e(b)/e(V)/r(table) triplet), this works for ANY command that
    populates r()/e() results - `summarize`, `tabstat`, `ci`, `svy: mean`,
    or a full e-class regression - you just name what you want back.

    Parameters
    ----------
    df : one replicate/bootstrap draw's data (supplied by
        StatCalculator.from_function per replicate weight).
    command : the Stata command to run, with "{weight}" as a placeholder
        for the weight column name if `weight` is given, e.g.
        "regress y x1 x2 [pw={weight}]" or "summarize x [aw={weight}]".
        Stata's weight syntax/placement varies enough by command (pw/aw/
        fw/iw, and where the bracket goes) that this doesn't try to build
        it for you - write it the way you'd type it directly in Stata.
        Pass a list of commands run in order, instead of a single string,
        when you need setup (a `svyset`, say) between `use` and the
        command whose results you actually want back - each element may
        use "{weight}", e.g. `["svyset psu [pw={weight}]", "svy: mean x"]`.
    results : names to pull back, e.g. ["r(mean)", "r(Var)"] or ["e(b)"].
        A scalar becomes one row (`join_on_name`=name); a matrix's columns
        (e.g. e(b)'s term names) each become their own row.
    weight : column name to substitute into "{weight}" in `command`, or ""
        (the StatCalculator.from_function convention - see its
        `weight_argument_name`) if `command` needs no weight or already
        hardcodes one. Default is "".
    join_on_name : name of the term-identifier column in the output.
        Default is "Variable".
    value_name : name of the estimate column in the output. Default is
        "estimate".
    edition, stata_path : forwarded to
        `_stata_interop.require_pystata`.
    reuse_data : if True, skip re-exporting/re-`use`-ing df when it's the
        same object (by identity) as the previous call - this is exactly
        the common case here: StatCalculator.from_function calls this
        delegate once per replicate weight with the *same* df object every
        time (only `weight`, spliced into `command`, changes), so the
        underlying data doesn't need re-exporting on every replicate.
        Default is False (safe/explicit opt-in - see
        `_stata_interop._run_in_stata`'s docstring for why this can't be a
        default the way the R side's dataframe_to_r caching is). Call
        `survey_kit.statistics._stata_interop.clear_stata_cache()` once
        the replicate loop is done, to free Stata's own copy of the data.
    quietly : pass False to let `command` stream Stata's own console
        output on success too (failures always surface the real error
        text regardless - see `_stata_interop._run_in_stata`'s
        docstring). Default is True.

    Returns
    -------
    pl.DataFrame
        One row per named scalar, or per column of a named matrix -
        exactly the shape a StatCalculator.from_function delegate needs
        (no vcov/tidy - see the note above on where the variance actually
        comes from).
    """
    from . import _stata_interop as _st

    if weight:
        if isinstance(command, str):
            full_command = command.format(weight=weight)
        else:
            full_command = [c.format(weight=weight) for c in command]
    else:
        full_command = command

    raw = _st.run_stata_results(
        df,
        full_command,
        results,
        reuse_data=reuse_data,
        edition=edition,
        stata_path=stata_path,
        quietly=quietly,
    )

    import numpy as np

    rows = []
    for name, value in raw.items():
        if isinstance(value, tuple):
            values, row_names, col_names = value
            arr = np.asarray(values)
            if arr.ndim == 1 or 1 in arr.shape:
                for term, v in zip(col_names, arr.reshape(-1)):
                    rows.append({join_on_name: term, value_name: float(v)})
            else:
                for i, rowi in enumerate(row_names):
                    for j, colj in enumerate(col_names):
                        rows.append(
                            {
                                join_on_name: f"{name}[{rowi},{colj}]",
                                value_name: float(arr[i, j]),
                            }
                        )
        else:
            rows.append({join_on_name: name, value_name: float(value)})

    return pl.DataFrame(rows)