NMAR
An R package for finite-population estimation when the probability of observing an outcome can depend on the outcome itself. NMAR puts several estimators behind one public interface and carries the same workflow across ordinary data frames and complex survey designs.
Why the missingness mechanism matters
With nonignorable nonresponse, missingness is part of the statistical model. A person's probability of responding may depend on the value that is missing, even after conditioning on observed covariates. In that setting, the respondent mean can be systematically different from the population mean.
NMAR targets finite-population means under explicit assumptions about that response process. The package does not hide those assumptions behind a generic missing-data switch. The estimator, formula structure and variance strategy remain visible in the engine configuration.
One interface, several estimators
The public call is nmar(formula, data, engine). The engine chooses the estimation method and owns its tuning parameters and formula interpretation. The current package includes empirical likelihood from Qin, Leung and Shao, parametric exponential tilting from Riddles, Kim and Im, and a fully nonparametric exponential-tilting method for aggregated categorical data.
This separation is useful beyond API neatness. New methods can reuse validation, scaling, bootstrap and result infrastructure without forcing users into another unrelated workflow. The package accepts both data.frame and survey.design inputs, with method-specific support and assumptions handled by the selected engine.
A common result contract
Statistical software becomes difficult to compose when every estimator invents its own output shape. NMAR uses a parent nmar_result abstraction with engine-specific subclasses. Fits can carry the primary estimate, standard error, convergence state, model coefficients, weights, sample information, inference metadata and numerical diagnostics in a predictable structure.
That shared contract powers familiar R methods such as summary(), se(), confint(), weights(), coef(), fitted(), tidy() and glance(). Engines can still expose method-specific details without making the basic inspection workflow change from one estimator to another.
Numerical reliability is part of the method
The hard part is not translating equations line by line into R. These estimators solve nonlinear systems, depend on fitted response probabilities and can be sensitive to scaling, starting values and poorly conditioned intermediate quantities.
The empirical-likelihood engine standardizes predictors before solving, optimizes the population response rate on a logit scale, guards extreme linear predictors and small denominators, exposes solver controls and reports convergence diagnostics. Bootstrap variance estimation is shared infrastructure rather than duplicated independently inside every engine, including support for survey-design resampling.
Same workflow, different engines
The estimator changes through the engine object rather than through a separate top-level API. The simplest comparison looks like this:
library(NMAR)
data("riddles_case1", package = "NMAR")
fit_el = nmar(
y ~ x,
data = riddles_case1,
engine = el_engine(variance_method = "none")
)
fit_et = nmar(
y ~ x,
data = riddles_case1,
engine = exptilt_engine(
y_dens = "normal",
family = "logit",
variance_method = "none"
)
)
summary(fit_el)
summary(fit_et)
The shared formula is only the surface. Each engine documents its own identification assumptions, formula interpretation, solver controls and variance options.
Research context
NMAR was developed within NCN OPUS 20 grant 2020/39/B/HS4/00941. The software formed the basis of my BEng thesis at Warsaw University of Technology and was presented at uRos 2025 and ElementsX 2025.
- Qin, Leung and Shao (2002): estimation with survey data under nonignorable nonresponse or informative sampling.
- Riddles, Kim and Im (2016): propensity-score adjustment for nonignorable nonresponse.
- NCN OPUS 20 research project, grant 2020/39/B/HS4/00941.