In [1]:
from __future__ import annotations

from survey_kit import logger
from survey_kit.statistics.adapters import r_feols, mi_ses_from_r_fixest
from sample_data import make_implicates
In [2]:
logger.info("The simplest way to run an R regression from survey_kit: r_feols(),")
logger.info("a wrapper around fixest::feols() with robust SEs by default. Requires R")
logger.info("itself plus rpy2/rpy2-arrow (`pip install survey-kit[r]`) and the R")
logger.info("'fixest' package. Check your setup cheaply with:")
logger.info("    from survey_kit.statistics._r_interop import check_r_setup")
logger.info("    check_r_setup(['fixest'])")
The simplest way to run an R regression from survey_kit: r_feols(),
a wrapper around fixest::feols() with robust SEs by default. Requires R
itself plus rpy2/rpy2-arrow (`pip install survey-kit[r]`) and the R
'fixest' package. Check your setup cheaply with:
    from survey_kit.statistics._r_interop import check_r_setup
    check_r_setup(['fixest'])
In [3]:
df_implicates = make_implicates()
logger.info(f"\n\nSample data: {len(df_implicates)} implicates, {df_implicates[0].height}")
logger.info("rows each (y = 1 + 2*x1 - 1.5*x2 + noise) - see sample_data.py.")

Sample data: 5 implicates, 300
rows each (y = 1 + 2*x1 - 1.5*x2 + noise) - see sample_data.py.
In [4]:
logger.info("\n\nOn one dataset, standalone - no MI at all:")
(df_estimates, df_ses, df_vcov, df_tidy) = r_feols(df_implicates[0], formula="y ~ x1 + x2")
logger.info(df_estimates)
logger.info(df_ses)

On one dataset, standalone - no MI at all:
shape: (3, 2)
┌─────────────┬───────────┐
│ Variable    ┆ estimate  │
│ ---         ┆ ---       │
│ str         ┆ f64       │
╞═════════════╪═══════════╡
│ (Intercept) ┆ 0.990009  │
│ x1          ┆ 2.016691  │
│ x2          ┆ -1.478573 │
└─────────────┴───────────┘
shape: (3, 2)
┌─────────────┬──────────┐
│ Variable    ┆ estimate │
│ ---         ┆ ---      │
│ str         ┆ f64      │
╞═════════════╪══════════╡
│ (Intercept) ┆ 0.02144  │
│ x1          ┆ 0.019448 │
│ x2          ┆ 0.024854 │
└─────────────┴──────────┘
In [5]:
logger.info("\n\nAcross multiple imputed datasets, combined via Rubin's rules -")
logger.info("mi_ses_from_r_fixest.feols(...) runs r_feols once per implicate and")
logger.info("combines the results, taking r_feols's own arguments directly:")
mi_reg = mi_ses_from_r_fixest.feols(
    df_implicates=df_implicates,
    formula="y ~ x1 + x2",
    round_output=False,
)
mi_reg.print(round_output=False)

Across multiple imputed datasets, combined via Rubin's rules -
mi_ses_from_r_fixest.feols(...) runs r_feols once per implicate and
combines the results, taking r_feols's own arguments directly:
Implicate #1
Implicate #2
Implicate #3
Implicate #4
Implicate #5
┌─────────────┬───────────┐
│    Variable ┆  estimate │
╞═════════════╪═══════════╡
│ (Intercept) ┆  0.983256 │
│             ┆  0.028668 │
│          x1 ┆  1.996435 │
│             ┆  0.047854 │
│          x2 ┆ -1.488951 │
│             ┆  0.030649 │
└─────────────┴───────────┘
In [6]:
logger.info("\n\nThat's it for the common case. mi_ses_from_r_fixest also has")
logger.info(".feglm/.fepois/.femlm for fixest's other estimators, all with the same")
logger.info("shape. r_lm_adapter (base R lm()/glm()) and r_fixest_adapter (any")
logger.info("fixest estimator via a func= string, including ones .feglm/.fepois/")
logger.info(".femlm don't cover) are the lower-level pieces these are built from -")
logger.info("see r_arbitrary_estimators.py for rolling your own with those, plus a")
logger.info("generic escape hatch into any R package/function at all.")

That's it for the common case. mi_ses_from_r_fixest also has
.feglm/.fepois/.femlm for fixest's other estimators, all with the same
shape. r_lm_adapter (base R lm()/glm()) and r_fixest_adapter (any
fixest estimator via a func= string, including ones .feglm/.fepois/
.femlm don't cover) are the lower-level pieces these are built from -
see r_arbitrary_estimators.py for rolling your own with those, plus a
generic escape hatch into any R package/function at all.