rchindia

Evidence-based maternal health insights across India

DLHS missing data imputation methods: choosing the right approach for survey analysis

A district-level reproductive-health estimate can move for reasons that have nothing to do with health.

UpdatedAugust 02, 2026
Read time14 min read
DLHS missing data imputation methods: choosing the right approach for survey analysis

A cluster of ages reported at 30 or 35, an unanswered question treated as a zero, or a skipped maternity-care item treated as a blank can alter a fertility or coverage measure before the analyst has fitted a single substantive model.

That is the hard part of DLHS missing data imputation methods in India: missingness is not one technical nuisance with one technical cure. It sits alongside questionnaire routing, age misreporting, eligibility screening, unequal selection probabilities, and clustered fieldwork. A plausible-looking completed dataset can still carry a distorted district story.

The right question is therefore not, “Which imputation command should we run?” It is: what kind of absence does this cell represent, what information survives around it, and what did the sampling design make visible—or leave outside the file altogether?

Accounting for DLHS Multistage Sampling Design in Imputation Models

DLHS-3 was not assembled as a simple random sample of Indian households. Its rural component used a two-stage stratified design; its urban component used three-stage stratification. Primary sampling units came from the 2001 Census frame, and household targets varied by district performance category defined from earlier RCH indicators.

That architecture matters twice. It matters when estimates are produced, because weights, strata, and clusters affect precision. It also matters earlier, while missing values are being handled, because the people available to “donate” information to a blank cell are not exchangeable observations in a flat spreadsheet.

A woman’s reported antenatal-care history is related to her schooling, parity, household circumstances, and district context. It is also related to the fieldwork environment around her: local service availability, interviewer conditions, social norms, and the PSU in which she was sampled. An imputation model that pools every respondent as if those contexts do not exist may smooth away precisely the variation a district analysis is supposed to detect.

In DLHS work, the sampling design is not a standard-error attachment added at the end. It is part of the data-generating story.

What design-aware work looks like

There is no single software setting that makes an imputation “survey-correct.” The practical aim is more modest and more demanding: preserve the relationships created by stratification, clustering, and unequal representation as far as the data and the model allow.

That usually means examining the following before selecting an imputation strategy:

  • Weights: Determine whether the analysis depends on weighted population inference, domain estimates, or an association model. The role of weights in imputation is not identical in each case. They may enter directly, inform predictors, or be incorporated through methods tailored to the eventual estimator.
  • PSU clustering: Where the model and available identifiers permit it, account for shared PSU-level context through multilevel specifications, cluster effects, or a carefully justified within-domain approach.
  • Strata and domains: Preserve meaningful strata and geographic domains in the model. Borrowing information across very different states or district types can manufacture homogeneity rather than recover a credible value.
  • Analytic estimand: A completed dataset for a district prevalence estimate is not automatically suitable for a regression, a fertility schedule, or a state comparison. The target analysis should shape the imputation model from the beginning.

The temptation is to fit a convenient regression, fill blanks with predictions, and apply survey weights later. That sequence can be defensible only in limited circumstances and after diagnostics. It is not a generic repair for DLHS microdata.

For district-level estimates, small effective sample sizes make this especially consequential. A completed value can affect both an estimate and the uncertainty around it. If the procedure treats imputed values as known facts, confidence intervals become artificially tidy—the statistical equivalent of a report that has been edited to remove all hesitation.

Addressing Age Heaping and Eligibility Boundary Errors Before Imputation

Age deserves separate treatment because two quite different problems are often collapsed under the label “missing data.” One concerns a recorded age that is likely rounded or misreported. The other concerns a woman who never entered the eligible interview universe because the screening age placed her outside it. The first may call for editing and, in narrow cases, model-based treatment. The second is undercoverage. There is no row, no response pattern, and no cell to fill.

Age heaping is a measurement problem, not a blank-value problem

The DLHS-3 quality study used Myers’ blended index to assess age heaping. Reported heaping was stronger in rural than urban areas, stronger among illiterate than literate respondents, and stronger in poorer than richer households. Those gradients matter because they overlap with the populations most central to fertility, maternal-care, and contraceptive analyses.

A value ending in zero or five is not automatically wrong. Nor should every heaped distribution trigger an attempt to “correct” individual ages. The danger lies in pretending that a generic regression can reverse a reporting process it does not observe. Replacing a heaped age with a model prediction based on education, wealth, and parity can make the age distribution look smoother while severing it from the respondent’s actual reported record.

For age-dependent indicators, start with an audit:

1. Inspect terminal-digit distributions by state, rural–urban residence, education, and other relevant subgroups.

2. Compare reported age against internally related fields, including marriage timing, parity, and birth-history dates where available.

3. Identify values that are impossible or inconsistent under the questionnaire logic.

4. Decide whether the problem is a limited set of invalid records, wider digit preference, or an eligibility-screening defect.

5. Keep the intervention proportional to the diagnosis. Editing an impossible value is different from redistributing an entire age profile.

The purpose is not to manufacture a perfectly smooth demographic curve. Real populations do not owe analysts smoothness. The purpose is to prevent clearly inconsistent records from being mistaken for usable predictors in later imputation.

Eligibility-boundary exclusion is not item nonresponse

DLHS individual interviews use defined age windows. The quality study documented exclusion around both lower and upper eligibility boundaries in multiple states, including substantial upper-boundary exclusion in some places. These women were screened out before the relevant individual interview. Their contraceptive histories, pregnancies, and care experiences are not latent answers waiting in the dataset.

A skipped question may be a missing response. A respondent excluded before interview is a coverage problem.

This distinction is decisive for handling missing data in demographic surveys. If older eligible women were systematically recorded as outside the age range, an analyst cannot repair the resulting bias by imputing values into the women who remain. Cell-level imputation operates on observed records. Eligibility exclusion changes the composition of the observed records themselves.

Appropriate responses may include sensitivity analyses, indicator bounds, revised weighting strategies where defensible, or explicit limitations in the reported estimate. The exact choice depends on the target indicator and what can be documented from the fieldwork and data structure. What it should not involve is converting an undercoverage issue into an item-nonresponse rate.

Evaluating Item Nonresponse Patterns and Auxiliary Variable Selection

Only after design and eligibility have been examined should the analyst ask what is genuinely missing at item level. This is where district-level household survey data cleaning often goes wrong: all blank-looking values are treated as one category.

They are not.

A blank can mean several different things

Questionnaire routing is part of the data. A woman who was not asked about antenatal care because she had no eligible birth does not have “missing ANC visits.” The item is not applicable. Coding it as a blank and then imputing a count creates a false numerator and a false denominator at once.

A useful classification is:

Recorded statusWhat it representsTypical treatment
Legitimate skip or not applicableThe questionnaire correctly bypassed the itemRetain as structurally inapplicable; do not impute
Refusal or unanswered reached itemA genuine item nonresponseInvestigate predictors and consider an imputation strategy
“Don’t know”A substantive inability or unwillingness to reportTreat separately from refusal where the mechanism may differ
Out-of-range or inconsistent responseA recorded value failing an edit ruleResolve through documented editing before considering imputation
Not interviewed because of eligibility or coverage failureMissing record, not missing itemAddress through coverage or sensitivity analysis, not cell imputation

This classification must be tied to the questionnaire and the coding documentation, not inferred casually from a single numeric missing-value code. One extract may combine several types of absence behind the same symbol. Another may retain route codes that disappear during cleaning. The first task is to preserve that provenance.

It is entirely possible for the raw share of blank-coded cells to look high and then fall sharply once legitimate skips are separated from genuine unanswered items. But no illustrative percentage should be treated as a property of DLHS-3 without computing it for the specific variable, population, and extract in hand. Missingness has a geography, a questionnaire location, and often a social pattern. It is not a universal rate printed on the dataset.

Build a missingness map before building a model

For each analytic variable, tabulate response status across the groups that matter for the proposed estimate: state, district where feasible, rural–urban residence, education, wealth position, age group, parity, and relevant service-use markers. Then inspect whether missingness is associated with the outcome’s plausible predictors and with the survey design variables.

This does not prove that missingness is missing at random. No logistic model of response can prove that. It does reveal whether a complete-case analysis is quietly removing a recognizable segment of the population.

The usual labels—MCAR, MAR, and MNAR—are useful as reasoning tools, not as badges of certainty.

  • MCAR is a strong claim and is rarely a safe working assumption for sensitive or recall-heavy RCH questions.
  • MAR can be a workable conditional assumption when rich observed predictors explain differences in response, but it depends on whether those predictors were actually retained and measured well.
  • MNAR remains possible when nonresponse depends on the unreported value itself, as may happen with stigmatized experiences, income-related questions, or reproductive events affected by disclosure concerns.

A model can be useful under MAR without making MNAR disappear. That is why sensitivity analysis belongs in the analysis plan rather than in the footnotes.

Auxiliary variables should do real work

The most useful auxiliaries are not simply variables with low missingness. They should have a credible relationship to both the item being completed and the probability that it is missing.

For reproductive and child health analyses, candidate auxiliaries may include education, household circumstances, parity, marital history, birth-history fields, residence, geographic identifiers, and related service-use variables. Their value depends on the target. Parity may help with a fertility-history item; it may be a poor stand-in for a question about a specific service experience. A district identifier may preserve geographic context but cannot substitute for individual-level predictors.

The imputation model should also be congenial with the final analysis. If the published model will relate institutional delivery to education, residence, and wealth, those relationships should not be absent from the completion model merely because a shorter model is easier to run.

There is a further practical constraint: an auxiliary variable that is itself heavily incomplete can spread instability through chained models. Sometimes it should be imputed jointly; sometimes it should be excluded; sometimes the target outcome should not be repaired at all. There is no prize for building the longest predictor matrix.

Implementing Design-Aware Multiple Imputation for Demographic Indicators

Multiple imputation can be a strong option for item nonresponse in DLHS analysis. It is not, however, an automatic DLHS-3 default, and chained equations are not a ritual requirement.

The choice should follow variable-specific diagnostics: the proportion and pattern of genuine nonresponse; the plausibility of a conditional missing-at-random assumption; the availability of strong auxiliaries; the outcome type; the number of observations within relevant domains; and the way weights, strata, and clustering enter the final estimator.

Where those conditions are reasonably supportive, a design-aware multiple-imputation approach can be preferable to complete-case analysis or single-value replacement because it carries uncertainty from the missing-data stage into inference. Where they are not supportive, it may merely give a sophisticated appearance to an unsupported assumption.

Choosing a method by the variable, not by habit

For a binary item with modest, plausibly explainable nonresponse and strong observed predictors, a logistic imputation model may be appropriate. For an ordered care-use measure, an ordinal approach may preserve the outcome structure better than treating categories as distances on a ruler. For continuous measures with irregular distributions, predictive mean matching can prevent implausible values, provided donors are selected within a model that respects relevant design and domain structure.

Chained equations are particularly useful when several variables have different forms and incomplete patterns. But they need careful specification:

  • Use a model compatible with each variable’s measurement scale and substantive range.
  • Include outcome relationships expected in the final analysis, not only generic demographics.
  • Incorporate strata, PSU-level structure, or multilevel terms when the identifiers and sample support them.
  • Avoid imputing across implausibly broad populations merely to increase the donor pool.
  • Check completed-data distributions against observed-data distributions within meaningful subgroups, rather than celebrating a smooth national average.
  • Generate enough completed datasets for estimates and uncertainty to stabilize; the appropriate number depends on the fraction of missing information and the intended analysis, not on an inherited software default.
  • Combine results using rules that reflect within- and between-imputation uncertainty, while retaining the complex-survey estimation procedure at the analysis stage.

Single imputation remains tempting because it yields one neat file. Mean substitution, deterministic regression fills, and one-pass hot-deck procedures can all narrow apparent variance and strengthen associations in misleading ways. They may have limited operational uses in descriptive data processing, but they should not be presented as uncertainty-aware solutions for district comparisons, regression coefficients, or demographic rates.

Diagnostics are the argument, not decoration

A credible imputation report should let a reader reconstruct the decision. It should state which response codes were treated as legitimate skips, which values were edited, what predictors entered the model, how survey design features were handled, and how results changed under plausible alternatives.

The most revealing comparison is often not between two algorithms. It is between analytic conclusions:

  • Do weighted district estimates materially change under complete-case and imputed analyses?
  • Are shifts concentrated in particular states, residence groups, or age bands?
  • Do imputed values preserve plausible relationships with parity, education, and service use?
  • Does the uncertainty widen appropriately once imputation is accounted for?
  • Does a result depend on a subgroup with too little observed information to support model-based completion?

If the answer to the last question is yes, restraint is better than false precision. An unstable district estimate should be described as unstable. A weakly observed variable may belong in a sensitivity appendix or a restricted analysis, not in the headline indicator.

Documentation keeps repair from becoming concealment

Every imputed value should remain traceable. Preserve the original response, the cleaned response status, an indicator showing whether the analytic value was imputed, and enough metadata to identify the method or model version used. This is not bureaucratic ornament. It allows later analysts to distinguish observed patterns from model-assisted ones.

Report response patterns in weighted terms where the target analysis is weighted. Explain how structural skips were handled. Identify any populations excluded by eligibility or coverage limitations. And do not imply that a completed dataset has “fixed” nonresponse. It has made a conditional statistical adjustment.

That language matters in health informatics as much as in survey methodology. A district estimate may travel from an analyst’s file into planning meetings, presentations, and resource decisions. Its uncertainty should travel with it.

The Discipline Behind a Defensible DLHS Estimate

The best DLHS dataset imputation techniques are rarely the most elaborate ones. They are the methods used after the analyst has separated a routing skip from a refusal, a heaped age from an impossible age, and an absent eligible respondent from a missing cell.

Start with the design. Audit age and eligibility boundaries. Map genuine item nonresponse. Select auxiliaries because they explain something, not because they are available. Then choose among complete-case analysis, weighting adjustments, multiple imputation, restricted-domain analysis, or sensitivity bounds according to the variable and the estimand.

For some DLHS indicators, design-aware multiple imputation will be the most credible route. For others, it will be unnecessary, weakly supported, or simply incapable of addressing the real defect. That is not a failure of the method. It is the point of doing the diagnosis first.

A completed cell is not automatically better data. In population analytics, the more useful achievement is a number whose limits are visible: what was observed, what was edited, what was modeled, and what remains unknown.

FAQ

Why is the multistage sampling design important during imputation?
The multistage design matters because respondents available to donate information to a blank cell are not exchangeable observations. Failing to account for stratification, clustering, and unequal representation can smooth away the exact district-level variation the analysis aims to detect.
What is the difference between age heaping and a blank-value problem?
Age heaping is a measurement issue involving rounded or misreported ages rather than a missing value. Attempting to correct heaped ages with a generic regression can sever the data from the respondent's actual reported record.
How should legitimate skips in the questionnaire be handled?
Legitimate skips or inapplicable items should be retained as structurally inapplicable and should not be imputed. Treating them as missing values creates false numerators and denominators.
What are the common classifications of missing status for questionnaire items?
Common classifications include legitimate skips, genuine item nonresponses from refusals, "don't know" responses indicating substantive inability to report, out-of-range or inconsistent responses, and exclusions due to eligibility or coverage failure.
When is multiple imputation appropriate for demographic indicators?
Multiple imputation is appropriate when there is genuine, explainable nonresponse, a plausible conditional missing-at-random assumption, strong auxiliary variables, and proper accounting for survey weights, strata, and clustering.