In [1]:
from __future__ import annotations
from survey_kit import logger
from survey_kit.statistics.adapters import r_feols, mi_ses_from_r_fixest
from sample_data import make_implicates
In [2]:
logger.info("The simplest way to run an R regression from survey_kit: r_feols(),")
logger.info("a wrapper around fixest::feols() with robust SEs by default. Requires R")
logger.info("itself plus rpy2/rpy2-arrow (`pip install survey-kit[r]`) and the R")
logger.info("'fixest' package. Check your setup cheaply with:")
logger.info(" from survey_kit.statistics._r_interop import check_r_setup")
logger.info(" check_r_setup(['fixest'])")
The simplest way to run an R regression from survey_kit: r_feols(),
a wrapper around fixest::feols() with robust SEs by default. Requires R
itself plus rpy2/rpy2-arrow (`pip install survey-kit[r]`) and the R
'fixest' package. Check your setup cheaply with:
from survey_kit.statistics._r_interop import check_r_setup
check_r_setup(['fixest'])
In [3]:
df_implicates = make_implicates()
logger.info(f"\n\nSample data: {len(df_implicates)} implicates, {df_implicates[0].height}")
logger.info("rows each (y = 1 + 2*x1 - 1.5*x2 + noise) - see sample_data.py.")
Sample data: 5 implicates, 300
rows each (y = 1 + 2*x1 - 1.5*x2 + noise) - see sample_data.py.
In [4]:
logger.info("\n\nOn one dataset, standalone - no MI at all:")
(df_estimates, df_ses, df_vcov, df_tidy) = r_feols(df_implicates[0], formula="y ~ x1 + x2")
logger.info(df_estimates)
logger.info(df_ses)
On one dataset, standalone - no MI at all:
shape: (3, 2) ┌─────────────┬───────────┐ │ Variable ┆ estimate │ │ --- ┆ --- │ │ str ┆ f64 │ ╞═════════════╪═══════════╡ │ (Intercept) ┆ 0.990009 │ │ x1 ┆ 2.016691 │ │ x2 ┆ -1.478573 │ └─────────────┴───────────┘
shape: (3, 2) ┌─────────────┬──────────┐ │ Variable ┆ estimate │ │ --- ┆ --- │ │ str ┆ f64 │ ╞═════════════╪══════════╡ │ (Intercept) ┆ 0.02144 │ │ x1 ┆ 0.019448 │ │ x2 ┆ 0.024854 │ └─────────────┴──────────┘
In [5]:
logger.info("\n\nAcross multiple imputed datasets, combined via Rubin's rules -")
logger.info("mi_ses_from_r_fixest.feols(...) runs r_feols once per implicate and")
logger.info("combines the results, taking r_feols's own arguments directly:")
mi_reg = mi_ses_from_r_fixest.feols(
df_implicates=df_implicates,
formula="y ~ x1 + x2",
round_output=False,
)
mi_reg.print(round_output=False)
Across multiple imputed datasets, combined via Rubin's rules -
mi_ses_from_r_fixest.feols(...) runs r_feols once per implicate and
combines the results, taking r_feols's own arguments directly:
Implicate #1
Implicate #2
Implicate #3
Implicate #4
Implicate #5
┌─────────────┬───────────┐ │ Variable ┆ estimate │ ╞═════════════╪═══════════╡ │ (Intercept) ┆ 0.983256 │ │ ┆ 0.028668 │ │ x1 ┆ 1.996435 │ │ ┆ 0.047854 │ │ x2 ┆ -1.488951 │ │ ┆ 0.030649 │ └─────────────┴───────────┘
In [6]:
logger.info("\n\nThat's it for the common case. mi_ses_from_r_fixest also has")
logger.info(".feglm/.fepois/.femlm for fixest's other estimators, all with the same")
logger.info("shape. r_lm_adapter (base R lm()/glm()) and r_fixest_adapter (any")
logger.info("fixest estimator via a func= string, including ones .feglm/.fepois/")
logger.info(".femlm don't cover) are the lower-level pieces these are built from -")
logger.info("see r_arbitrary_estimators.py for rolling your own with those, plus a")
logger.info("generic escape hatch into any R package/function at all.")
That's it for the common case. mi_ses_from_r_fixest also has
.feglm/.fepois/.femlm for fixest's other estimators, all with the same
shape. r_lm_adapter (base R lm()/glm()) and r_fixest_adapter (any
fixest estimator via a func= string, including ones .feglm/.fepois/
.femlm don't cover) are the lower-level pieces these are built from -
see r_arbitrary_estimators.py for rolling your own with those, plus a
generic escape hatch into any R package/function at all.