rchindia

Evidence-based maternal health insights across India

DLHS and NFHS data discrepancies: how to reconcile them

The problem with comparing DLHS and NFHS data is not that one survey is “right” and the other is “wrong.” The problem is that they were built to measure overlapping parts of India’s health system under different operating conditions.

UpdatedAugust 04, 2026
Read time22 min read
DLHS and NFHS data discrepancies: how to reconcile them

Change the sampled population, the reference period, the geography, or the age band, and the apparent trend can move even when the underlying health condition has not changed.

That is the core of DLHS–NFHS data discrepancy reconciliation. Researchers are not simply joining two files. They are repairing breaks in the measurement system.

DLHS was designed to deliver detailed district-level information on reproductive and child health. NFHS operated as a broader national health and population survey, with a different architecture and, in earlier rounds, a different balance between national, state, and district estimates. Starting with NFHS-4 in 2015–16, the Ministry of Health and Family Welfare subsumed DLHS and the Annual Health Survey into NFHS. The change reduced duplication, but it did not erase the comparability problems in the historical record.

A clean-looking trend line can therefore be a statistical failure disguised as a result.

Structural evolution: from parallel surveys to integrated NFHS-4

The first task in comparing DLHS and NFHS datasets is to establish what each round was actually built to do.

The survey sequence is not a single uninterrupted measurement program:

  • NFHS-1 was conducted in 1992–93.
  • NFHS-2 and DLHS-1 were conducted in 1998–99.
  • DLHS-2 followed in 2002–04.
  • NFHS-3 was conducted in 2005–06.
  • DLHS-3 followed in 2007–08.
  • DLHS-4 was conducted in 2012–13.
  • NFHS-4 was conducted in 2015–16.
  • NFHS-5 was conducted in 2019–21.

These rounds overlap in subject matter, but not necessarily in sample design, field implementation, respondent eligibility, laboratory or anthropometric protocols, or geographic coverage.

The scale of NFHS changed sharply at NFHS-4. NFHS-3 included 109,041 households. NFHS-4 expanded to 601,509 households—roughly five times the earlier NFHS sample—and produced district-level estimates for the first time. That move brought NFHS into territory previously associated with DLHS: local-area estimates that could be used for district planning.

DLHS-3, by contrast, covered 720,320 households across 601 districts. Its operating logic was more explicitly district-oriented. The survey was built to provide granular information for reproductive and child health programs, service utilization, maternal care, immunization, contraception, and related indicators.

That distinction matters because sample size is not the only determinant of usefulness. A larger survey may support district estimates, but it does not automatically become interchangeable with an older district survey. The instrument, field period, eligibility rules, weighting, and geographic frame still determine what the estimate means.

Integration solved an administrative duplication problem. It did not magically make every historical indicator comparable.

NFHS-4’s expansion was a major improvement for national and district-level coverage, but it also created a temptation: place the NFHS-4 estimate beside the closest DLHS estimate and call the difference a trend. That shortcut fails whenever the two estimates are based on different populations or definitions.

The operational break in 2015–16

The decision to subsume DLHS and AHS into NFHS-4 created a common platform for subsequent national health and population measurement. It also established NFHS as the main instrument for district-level estimates at scale.

For analysts, the practical implication is straightforward:

1. Treat pre-NFHS-4 DLHS and NFHS rounds as related but distinct measurement systems.

2. Treat NFHS-4 as a methodological transition point, not merely another observation in a time series.

3. Document the exact indicator definition before calculating change.

4. Use harmonization methods when combining rounds rather than assuming that shared labels imply shared measurement.

The word “stunting,” for example, is not enough. The analyst needs to know the child age range, the measurement protocol, the handling of missing values, and the survey population included in the denominator. Without that metadata, the number has no stable meaning.

The most common failure in demographic survey discrepancy analysis is a mismatch in denominators.

A survey indicator is not just a percentage. It is a percentage among a defined group, measured over a defined period, using a defined question or clinical protocol. Alter any of those components and the estimate changes.

Child nutrition is a direct example

DLHS-2 covered children under five years of age for relevant child nutrition measures. NFHS-3 used children under three years for comparable nutrition indicators. Those are not equivalent populations.

Children aged three and four are not a statistical footnote. Growth patterns, feeding practices, illness exposure, and survival selection differ across early childhood. An underweight or stunting estimate among children under three cannot be read as a direct continuation of an estimate among children under five.

The same problem appears in reproductive and maternal health indicators. A measure may refer to:

  • women aged 15–49;
  • currently married women;
  • all women who were married;
  • recent births within a specified reference period;
  • all births in the preceding five years;
  • the most recent birth only.

These groups overlap, but they are not interchangeable.

Biomarker eligibility changed as well

The clinical, anthropometric, and biochemical components of DLHS-4 and AHS included all non-pregnant participants aged 18 years or older. NFHS-4 used a different biomarker structure: women aged 15–49 years and a 15% subsample of men aged 15–54 years.

That difference affects both prevalence and representativeness. A biomarker estimate from adults aged 18 and above cannot be treated as equivalent to one focused on women aged 15–49 and a selected male subsample. The age distribution is different. The sex composition is different. The sampling fraction is different. The result is a different estimate of a related phenomenon.

If an analyst compares those figures without adjustment, the spreadsheet may calculate a percentage change, but the study has not established a valid temporal change.

Reference periods are another hidden fault line

Survey questions often depend on a recall window:

  • illness in the previous two weeks;
  • births in the previous five years;
  • contraceptive use at the time of interview;
  • antenatal care received during a recent pregnancy;
  • service contact during a specified period.

A survey fielded over one interval and another fielded over a different interval can produce different results because the reference periods capture different events, seasons, or policy environments. A two-week illness estimate can shift with monsoon conditions, local outbreaks, or seasonal access to care. A five-year birth history can conceal short-term changes in service delivery.

A WHO feasibility assessment of trends across NFHS, AHS, and DLHS found that differences in reference periods and reference groups made direct temporal comparisons difficult without adjusting the baseline. That is the correct interpretation: a discrepancy may be a measurement mismatch before it is a health-system signal.

Build a comparability matrix before merging anything

A practical reconciliation process begins with a variable-level matrix. For every candidate indicator, record the following:

  • survey round and field dates;
  • geographic unit and boundary definition;
  • eligible population;
  • numerator definition;
  • denominator definition;
  • recall or reference period;
  • measurement method;
  • weighting and adjustment variables;
  • missing-data treatment;
  • whether the value is a direct estimate or a model-based estimate.

A compact example:

ParameterDLHS-2 / DLHS-4 patternNFHS-3 / NFHS-4 patternReconciliation action
Child nutrition age bandDLHS-2 used children under fiveNFHS-3 used children under threeDo not compare as a single trend without recalculation or age-standardization
Biomarker populationDLHS-4 and AHS included non-pregnant participants aged 18+NFHS-4 tested women aged 15–49 and a 15% male subsample aged 15–54Restrict analysis to a genuinely overlapping population, where possible
GeographyDLHS rounds were district-oriented; coverage and boundaries variedNFHS-4 covered all states and Union Territories and supplied district estimatesHarmonize district boundaries and flag unmatched areas
Reference periodVaries by indicator and roundVaries by indicator and roundCompare only indicators with aligned recall windows
National coverageDLHS-4 excluded nine low-performing statesNFHS-4 covered all states and Union TerritoriesAvoid unadjusted national comparisons

This is not bureaucratic overhead. It is the minimum infrastructure required to prevent an invalid result.

Geographic coverage gaps: the missing states are not random

DLHS-4 excluded nine low-performing states, including Bihar, Madhya Pradesh, and Uttar Pradesh. Those states had historically higher levels of malnutrition and other health-system pressures. As a result, DLHS-4 does not provide a complete national picture for indicators affected by that exclusion.

This creates a classic coverage problem: the missing geography is connected to the outcome being measured.

If the omitted states had been average-performing areas, a national estimate might still be distorted, but the direction of bias would be harder to anticipate. When low-performing states are excluded, a DLHS-4 aggregate can look stronger than a full-coverage NFHS-4 estimate even if conditions did not improve between rounds.

That is not evidence of deterioration by itself. It may be a change in the population represented.

Three geographic questions must be answered

Before comparing district-level health data, establish:

1. Are the same districts present in both rounds?

District names are not enough. Administrative boundaries can change, districts can split, and coding systems can be revised.

2. Does the aggregate cover the same states?

A state-level or national average built from different state sets is not a like-for-like comparison.

3. Are the weights designed for the same estimand?

A district estimate, a state estimate, and a population-weighted national estimate answer different questions.

For district analysis, researchers may need to create a crosswalk between historical district boundaries and the boundaries used in later rounds. Where a district has been divided, the analyst cannot simply copy the old value into each successor district. That duplicates information and falsely increases apparent precision.

Where districts have merged, an old estimate may not map cleanly to the new administrative unit. The correct approach may be aggregation using population or household weights, provided the underlying microdata and the required auxiliary information are available. If they are not, the limitation needs to remain visible in the result.

Why national averages can conceal the break

Suppose one round excludes several high-risk states and another includes them. The national mean will shift because the composition of the national sample changed. State performance, population size, and survey weighting all enter the calculation.

A robust comparison should therefore show at least two views:

  • the full-coverage estimate for each survey round;
  • a restricted comparison using only the common geographic coverage.

The restricted comparison does not solve every problem, but it separates geographic composition effects from changes in the shared sample frame. If the trend appears in the common-coverage analysis and in the full-coverage analysis, confidence increases. If it appears only after adding previously excluded states, the analyst should describe a coverage-driven shift rather than a simple trend.

A national average is only as stable as the map underneath it. Change the map and the average has changed before the health outcome has.

Quantifying data quality: age reporting is a field problem, not a footnote

Age misreporting is common in demographic surveys. Respondents may not know their exact birth date. Interviewers may accept rounded ages. Household members may report someone else’s age from memory. Administrative records may be absent, inconsistent, or unavailable.

The result is age heaping: an unusual concentration of reported ages ending in particular digits, especially zero and five. That distortion matters because age is used to define eligibility, construct denominators, calculate fertility measures, interpret mortality, and compare child growth across age bands.

In the reported evidence, age reporting quality was poor across both NFHS and DLHS rounds. Myer’s Index values were 9.6 and 10.5 for NFHS-II and III, and 12.1 and 10.8 for DLHS-II and III. Whipple’s Index values were 1.9 and 2.1 for NFHS-II and III, compared with 2.4 and 2.1 for DLHS-II and III. The U.N. Joint Score was 28.4 and 30.5 for NFHS-II and III, and 23.8 and 29.1 for DLHS-II and III.

The exact index interpretation depends on the scale and method used, but the operational conclusion is not ambiguous: age reporting cannot be treated as perfectly accurate in either survey family.

What each index is telling you

  • Whipple’s Index focuses on preference for ages ending in selected digits, commonly five and zero. It is useful for detecting concentrated heaping.
  • Myer’s Blended Index examines the preference for terminal digits across a broader set of ages.
  • The U.N. Joint Score combines age and sex distribution issues into a broader diagnostic of population reporting quality.

These measures do not repair the dataset. They tell the analyst how much caution is required.

A high level of age heaping can affect comparisons in at least four ways:

1. It shifts respondents into or out of eligibility bands.

2. It changes the apparent age structure of women of reproductive age.

3. It distorts age-specific fertility or service-use rates.

4. It creates artificial differences between survey rounds if the reporting pattern changes.

Practical responses to age misreporting

The right correction depends on the indicator.

For broad age groups, analysts may use grouped categories that are less sensitive to single-year reporting errors. For age-specific analysis, smoothing or redistribution methods may be required. For child indicators, anthropometric analysis should follow the survey’s age calculation and flag implausible or uncertain ages rather than silently replacing them.

The analyst should also compare the age distribution across rounds before estimating change. Look for:

  • spikes at ages ending in zero or five;
  • implausible declines or increases at age boundaries;
  • sudden shifts in the share of women entering or leaving the 15–49 age range;
  • differences in the proportion of respondents with unknown or imputed ages.

If an indicator depends heavily on a narrow age band, age reporting quality becomes part of the substantive result. It cannot be relegated to a methods appendix.

Harmonizing the historical record without manufacturing precision

Once the population definitions and geographic frames are mapped, the next question is how to combine or compare the estimates.

There is no universal formula for merging historical DLHS and NFHS data. The exact mathematical procedure used by the Ministry of Health and Family Welfare to weight and merge earlier datasets before the formal subsuming in NFHS-4 is not established in the available evidence. Analysts should not imply that a single official conversion factor exists for every indicator.

Instead, reconciliation should be indicator-specific.

Start with the estimand

Before selecting a method, state exactly what is being estimated:

  • change in prevalence among children under a common age threshold;
  • change in contraceptive use among a defined group of women;
  • district-level service coverage for a common set of districts;
  • state-level variation after geographic standardization;
  • national prevalence under a fixed population composition.

If the estimand is vague, the harmonization will be vague too.

Use restricted comparisons when possible

The cleanest comparison often uses the overlapping population and geography rather than forcing every available record into one model.

For example:

  • restrict both rounds to a common age band;
  • restrict both rounds to the same sex and pregnancy status;
  • use common districts only;
  • align the reference period;
  • apply consistent exclusion rules for missing observations.

This reduces the amount of data, but it improves the meaning of the comparison. A smaller valid comparison is more useful than a national-looking estimate assembled from incompatible pieces.

Standardize before calculating change

When age structures differ between survey rounds, direct prevalence comparisons can be misleading. A standardization procedure applies a common reference age distribution to each round. The result answers a controlled question: what would prevalence have been in each survey if both populations had the same age composition?

The reference distribution can come from a census or another declared standard, but it must be documented. The choice of standard affects the result, particularly when age-specific rates vary substantially.

The same principle applies to geographic composition. If the goal is to compare health outcomes rather than population mix, use common-area or population-standardized estimates. Do not bury the composition adjustment inside an unexplained weighting step.

Treat survey weights as part of the measurement design

Survey weights compensate for unequal probabilities of selection, nonresponse, and sometimes post-stratification. They are not interchangeable across rounds.

A common error is to append survey files and apply one pooled weight as though all observations came from one sample design. That can produce a precise-looking estimate with no defensible sampling interpretation.

For separate-round comparisons:

  • calculate each estimate using its own survey design;
  • preserve strata and cluster information;
  • account for primary sampling units;
  • report uncertainty for each estimate;
  • test whether observed differences exceed expected sampling variation.

For pooled modeling:

  • include survey round as a design component;
  • use round-specific weights or a defensible rescaling strategy;
  • account for clustering and stratification;
  • specify how the combined target population is defined.

This is where many population analytics projects fail. The code runs. The table exports. The standard errors are wrong.

Small Area Estimation: repairing unstable district estimates

District-level analysis creates a second technical problem: sample sizes can be too small for reliable direct estimates.

A district estimate may be available in the file but still be unstable. Wide confidence intervals, sparse subgroups, high design effects, and missing clusters can make the estimate unsuitable for ranking districts or detecting small changes.

Researchers address this with Small Area Estimation (SAE). SAE combines survey observations with auxiliary information to produce model-based estimates for areas where the direct survey sample is limited.

In the Indian context, researchers have used auxiliary variables from the 2001 and 2011 Censuses alongside survey data. The basic logic is practical:

  • the survey contributes the health outcome;
  • the census contributes population and area-level predictors;
  • the model borrows strength across related districts;
  • the resulting estimate is designed to be more stable than the raw district proportion.

SAE is not a free precision upgrade. It changes the source of information and introduces model dependence.

What SAE can fix

SAE can help with:

  • small district-level sample sizes;
  • high sampling variability;
  • unstable subgroup estimates;
  • gaps where direct estimates are too noisy for planning;
  • consistent prediction across areas with related auxiliary characteristics.

It can also support historical comparisons when the survey instruments differ, provided the analyst models the differences explicitly rather than pretending they do not exist.

What SAE cannot fix on its own

SAE does not automatically correct:

  • incompatible age definitions;
  • different biomarker eligibility rules;
  • missing states in one survey;
  • changes in district boundaries;
  • inconsistent reference periods;
  • systematic measurement error;
  • poor outcome definitions;
  • unexamined age misreporting.

A model can smooth an unstable estimate. It cannot turn a measure of children under three into a measure of children under five. It cannot make an excluded state part of a survey’s observed sample. It cannot repair a denominator that changed between rounds.

A defensible SAE workflow

A field-ready SAE workflow should include:

1. Define the target area.

Fix the district boundary version and establish how historical districts map to current units.

2. Select auxiliary variables available across the required period.

Census indicators, population composition, infrastructure, education, and service-access variables may be useful, but availability and consistency matter more than model glamour.

3. Model the direct survey estimate and its uncertainty.

The model needs to know which districts have strong direct evidence and which do not.

4. Validate against areas with adequate sample sizes.

Compare model-based estimates with direct estimates and inspect systematic residuals.

5. Run sensitivity checks.

Change the auxiliary-variable set, specification, and geographic treatment. If rankings change substantially, the result is not robust.

6. Report model-based estimates as model-based.

Do not publish them in a table that makes them look like direct survey observations.

The point is not to produce the smoothest map. It is to produce an estimate that planners can use without confusing modeled stability with observed certainty.

A reconciliation protocol that survives scrutiny

A robust comparison of DLHS and NFHS data should move through the following sequence.

1. Freeze the indicator definition

Write the numerator and denominator in plain language. Include age, sex, pregnancy status, marital status, birth history, recall window, and measurement method.

If the definition cannot fit into one clear sentence, the indicator is not ready for cross-survey comparison.

2. Map the survey populations

Create a side-by-side record for every round. Note who was eligible, who was tested, who was interviewed, and who was excluded. Biomarker data deserve special treatment because the tested subsample may differ sharply from the general household sample.

3. Align geography

Build a district and state crosswalk. Flag:

  • renamed districts;
  • split districts;
  • merged districts;
  • missing districts;
  • states excluded from a round;
  • changes in urban and rural classification.

Do not resolve an unmatched geography by copying values or assigning them based on name similarity alone.

4. Check age quality

Run Whipple’s, Myer’s, or U.N. Joint Score diagnostics where the required age data are available. Plot the age distribution. Inspect eligibility thresholds. If age heaping is severe, adjust the analytic strategy or downgrade the strength of the conclusion.

5. Select the comparison design

Choose one of three defensible paths:

  • Direct restricted comparison: same age band, geography, and reference period.
  • Standardized comparison: estimates adjusted to a common age or geographic composition.
  • Model-based reconciliation: SAE or another explicit model used to address small areas or incomplete coverage.

State why the chosen design fits the research question.

6. Quantify uncertainty

Report confidence intervals or other uncertainty measures. For model-based estimates, report the model uncertainty as well as the survey component where appropriate. A two-percentage-point difference with wide overlapping uncertainty is not the same finding as a two-percentage-point difference measured precisely.

7. Run a sensitivity analysis

At minimum, compare:

  • full geographic coverage versus common geographic coverage;
  • unstandardized versus standardized estimates;
  • direct district estimates versus SAE estimates;
  • strict age-band restriction versus broader inclusion;
  • alternative treatments of missing or implausible ages.

If the conclusion changes under reasonable specifications, write that into the result. That is not weakness. It is the actual structure of the evidence.

8. Keep the break visible

NFHS-4 should be marked as a transition in the series. A chart that displays DLHS-2, DLHS-3, DLHS-4, and NFHS-4 as equally comparable points is visually convenient and analytically dangerous.

Use annotations. Separate panels if necessary. Explain the coverage and instrument changes directly beside the figure, where the reader will see them.

What researchers should not do

Several shortcuts recur in published and internal analyses:

  • comparing DLHS-2 child nutrition estimates under five with NFHS-3 estimates under three;
  • treating DLHS-4 as a complete national estimate despite its exclusion of nine low-performing states;
  • comparing biomarker prevalence across different age and sex eligibility rules;
  • joining district records by district name without boundary reconciliation;
  • interpreting every difference as a policy effect;
  • ignoring age heaping because the sample size is large;
  • pooling survey records without preserving each round’s design and weights;
  • reporting SAE outputs as if they were direct observations;
  • presenting a smooth time series without marking methodological breaks.

The common thread is operational: the analyst is optimizing for a complete table instead of a defensible measurement.

That trade-off is backwards. In population analytics, missingness and incompatibility should be visible. A blank cell with an explanation is better than a fabricated continuity.

The scalable solution: build a metadata layer, not another spreadsheet

The durable fix is to treat survey reconciliation as infrastructure.

Every indicator should have a metadata record containing:

  • source survey and round;
  • field period;
  • target population;
  • numerator and denominator;
  • geographic frame;
  • reference period;
  • instrument version;
  • measurement protocol;
  • weighting method;
  • age-quality diagnostics;
  • direct or model-based status;
  • comparability rating;
  • known exclusions and breaks.

This metadata layer should sit beside the analysis dataset, not in a forgotten methods document. When a district estimate is updated, the analyst should be able to see immediately whether the change reflects a new survey observation, a boundary revision, a revised weight, or a model update.

A useful comparability rating can be operational rather than decorative:

  • High comparability: common age definition, geography, reference period, and measurement method.
  • Moderate comparability: one component standardized or restricted, with residual limitations.
  • Low comparability: major differences in population, coverage, or method; descriptive use only.
  • Not comparable: no defensible bridge between the estimates.

This prevents a familiar failure in health information systems: a dashboard that displays every number with the same visual confidence.

The job is not to eliminate every discrepancy. The job is to identify which discrepancies are fixable, which are modelable, and which must remain warnings on the map.

DLHS and NFHS remain valuable precisely because they cover different periods of India’s reproductive, maternal, and child health system and provide evidence at different levels of geographic detail. But their value is lost when analysts force them into a false continuous series.

The correct route is disciplined and repeatable: define the estimand, align the population, reconcile the map, test age quality, preserve survey design, use standardization or SAE where justified, and publish the uncertainty. NFHS-4’s integration of district-level coverage was a major structural shift. It should be used as one—not hidden.

A credible analysis may produce fewer comparable points, wider intervals, and more methodological notes than a quick chart. That is acceptable. The alternative is a polished trend line built on incompatible parts, the statistical equivalent of repairing a broken supply chain with relabeled boxes.

FAQ

Why do DLHS and NFHS surveys show different health trends for the same indicators?
The differences occur because the surveys use different sampled populations, age bands, reference periods, and geographic coverage. Changing any of these components alters the estimate, creating apparent trends even when the underlying health condition has not changed.
Why can't I directly compare child nutrition data from DLHS-2 and NFHS-3?
DLHS-2 measured child nutrition for children under five years of age, while NFHS-3 used children under three years for comparable indicators. These are not equivalent populations, as growth patterns and illness exposure differ significantly between children aged three to four and younger children.
How does the exclusion of certain states in DLHS-4 affect national health estimates?
DLHS-4 excluded nine low-performing states with historically higher levels of malnutrition, which makes its national aggregate look artificially stronger than full-coverage surveys. This creates a coverage problem where the missing geography is directly connected to the health outcomes being measured.
What is the purpose of Small Area Estimation in district-level health analysis?
Small Area Estimation combines survey observations with auxiliary census information to produce more stable, model-based estimates for districts where the direct survey sample is too small or noisy. It helps repair unstable district estimates but cannot fix incompatible age definitions or missing states.
How does age misreporting and heaping impact demographic survey data?
Age misreporting causes age heaping, where reported ages concentrate on digits like zero and five, which shifts respondents into or out of eligibility bands and distorts age-specific rates. This distortion affects the calculation of fertility measures, mortality interpretation, and child growth comparisons across different age bands.