DLHS indicator definition shifts: trend analysis traps
A percentage point is not necessarily a population change. In the District Level Household Survey series, a difference between two published estimates may reflect a new respondent universe, altered…

A percentage point is not necessarily a population change. In the District Level Household Survey series, a difference between two published estimates may reflect a new respondent universe, altered geographic coverage, a different recall window, sampling error, or a label that masks a different numerator. The arithmetic is simple. The comparability is not.
DLHS-1 was conducted in 1998–99. DLHS-2 followed in 2002–04, DLHS-3 in 2007–08, and DLHS-4 in 2012–13. These rounds provide an unusually valuable record of reproductive and child health indicators across Indian districts. They are also repeated cross-sectional surveys. They did not follow the same women, households, or births over time. A line chart connecting four estimates can therefore imply a continuity that the survey design does not supply.
The central problem in DLHS trend analysis is usually described as “indicator definition shifts.” That description is incomplete. Some indicators did retain consistent definitions. The more persistent threat is denominator instability: who was eligible, where household data were collected, which events entered the reference period, and whether two similarly named measures are in fact measuring the same construct.
A trend is not created by placing estimates in chronological order. It is created by making their populations, exposure periods, and definitions commensurate.
The evolution of respondent eligibility
The most consequential discontinuity occurs between the earlier rounds and DLHS-3. DLHS-1 and DLHS-2 interviewed currently married women aged 15–44. DLHS-3 expanded the female respondent universe. It included ever-married women aged 15–49 and never-married women aged 15–24.
This is not a minor questionnaire revision. It changes the denominator for many reproductive health indicators.
A rate calculated among currently married women aged 15–44 answers a specific question: what was the measured experience, knowledge, or service use within that marital and age-defined population? An estimate among ever-married women aged 15–49 answers another. The latter includes women aged 45–49 and women who were formerly married. Those groups can have materially different fertility histories, contraceptive exposure, pregnancy experience, service use, and health-seeking profiles.
The effect varies by indicator.
For current contraceptive use, the inclusion of women outside the earlier currently married 15–44 universe may lower or raise a total estimate depending on the composition of the added population. For completed fertility or parity indicators, the addition of women aged 45–49 will predictably increase average lifetime exposure to childbearing. For maternal care indicators based on recent births, the direction is less automatic, but the change in eligible women still alters the base population from which births are reported.
DLHS-3 did publish some tabulations separately for women aged 15–49 and for currently married women aged 15–44. This distinction should govern any DLHS-2-to-DLHS-3 comparison. Where the same indicator and reference period are available, the DLHS-3 subgroup of currently married women aged 15–44 is the defensible counterpart to the DLHS-2 universe.
| Analytical component | DLHS-1 and DLHS-2 | DLHS-3 | Consequence for trend analysis |
|---|---|---|---|
| Core female respondent universe | Currently married women aged 15–44 | Ever-married women aged 15–49; never-married women aged 15–24 | Published all-women estimates are not automatically comparable |
| Upper age limit | 44 years | 49 years for ever-married women | Later-age reproductive histories enter some estimates |
| Marital-status scope | Currently married only | Includes formerly married women among ever-married respondents | Exposure, fertility history, and service-use profiles may change |
| Appropriate bridge for many indicators | Currently married women 15–44 | DLHS-3 tabulation for currently married women 15–44, where available | Reduces denominator discontinuity |
A common error is to take the published DLHS-2 figure and the published DLHS-3 figure from a national report table, calculate the difference, and label it improvement or decline. That calculation may be numerically correct and analytically invalid.
The correction is not to discard the data. It is to define the target population before extracting the estimate. If the target is currently married women aged 15–44, retain that universe in every comparable round. If the target is all ever-married women aged 15–49, an early-round equivalent may not exist. In that case, the trend cannot be reconstructed from headline estimates without qualification.
DLHS-4 is not a national household continuation
DLHS-4 introduces a separate break. Earlier DLHS rounds provided broad all-India household survey coverage. DLHS-4 did not replicate that geography for household indicators.
In DLHS-4, household and facility surveys were conducted in non-AHS states. In the nine AHS states, only facility surveys were conducted. The AHS states were Assam, Bihar, Chhattisgarh, Jharkhand, Madhya Pradesh, Odisha, Rajasthan, Uttar Pradesh, and Uttarakhand.
This distinction is decisive for national comparisons. Those states represent a large and demographically distinct segment of India’s population. Their exclusion from the DLHS-4 household universe means that an apparent national change may instead be a change in the map.
A household-based maternal health estimate from DLHS-4 cannot be treated as directly comparable with an earlier all-India DLHS household estimate unless the earlier data are restricted to the same non-AHS geography, or a valid harmonized approach explicitly accounts for the coverage difference. Adding the label “India” to both columns does not solve the problem.
The problem is especially acute for reproductive health indicators India analysts often use to describe regional inequality: institutional delivery, antenatal care, contraception, child immunisation, fertility-related measures, and household-reported morbidity. The AHS states had distinct demographic structures and health-system conditions. Excluding them can change the aggregate even when every district-level rate remains fixed.
A defensible DLHS-3-to-DLHS-4 exercise therefore begins with geography, not with the indicator table:
1. Identify whether the DLHS-4 measure is a household indicator or a facility indicator. Facility data and household data do not answer interchangeable questions. A facility readiness measure cannot be used as a direct continuation of household-reported service use.
2. Restrict the earlier survey round to the DLHS-4 household geography. For a non-AHS trend, the comparator must exclude the nine AHS states from DLHS-3 and earlier rounds.
3. Recalculate aggregate weights where microdata and survey design information permit. A simple mean of district percentages is not a population estimate. Districts differ sharply in population size and sampling design.
4. Label the result accurately. “Non-AHS-state trend” is an analytical description. “India trend” is not, unless the excluded geography has been accounted for.
5. Keep facility and household series separate. A facility survey may illuminate service capacity. It cannot repair the missing household denominator.
DLHS-4 household estimates describe a restricted geography. A national label does not restore the excluded population.
This is one of the most serious DLHS data comparison errors because it can reverse interpretation. Suppose a non-AHS aggregate shows a higher level of a maternal care indicator than the earlier national estimate. The difference may be real. It may also reflect removal of states with different baseline conditions. Without a matched-geography comparison, the data do not distinguish the two explanations.
Composite indicators are not their labels
Composite indicators create a second category of error. Analysts frequently substitute familiar labels for formal definitions. This is methodologically unsafe.
The distinction between “safe delivery” and “institutional delivery” is the clearest example. In DLHS-2 documentation, safe delivery is broader than institutional delivery. It includes a delivery in an institution or a home delivery assisted by a doctor, auxiliary nurse midwife, or nurse.
These indicators overlap. They are not equivalent.
Institutional delivery identifies place of delivery. Safe delivery incorporates an assistance criterion and includes some home births. A rise in institutional delivery may contribute to a rise in safe delivery, but the two series have different numerators. Relabelling one as the other creates an artificial trend even if the survey instrument has not changed.
Full antenatal care requires equal precision. Official DLHS-2 and DLHS-3 documentation defines full ANC using three components:
- at least three antenatal care visits;
- at least one tetanus toxoid injection; and
- 100 or more iron-folic-acid tablets, or equivalent syrup.
The available documentation does not support a claim that this composite definition changed between DLHS-2 and DLHS-3. A survey round changed. The formal three-part definition did not, on the evidence available.
That does not make every full ANC comparison automatically clean. The reference population may differ. The reference period may differ. Recall error can affect reported visits, injections, and tablet consumption. District-level confidence intervals may overlap even when point estimates do not. But these are distinct issues. They should not be obscured by an unsupported assertion that the indicator itself was redefined.
| Measure | What it captures | Frequent analytical error | Correct treatment |
|---|---|---|---|
| Institutional delivery | Births occurring in a health institution | Treating it as identical to safe delivery | Use only for place-of-delivery trends |
| Safe delivery | Institutional births plus eligible skilled-assisted home births | Relabelling as institutional delivery | Retain the broader definition in all reporting |
| Full ANC | Three ANC visits, tetanus toxoid, and 100+ IFA tablets or equivalent syrup | Assuming a definition shift merely because the survey round changed | Confirm the denominator and recall period, then compare like with like |
| ANC visit count | Contact frequency | Treating it as a complete quality-of-care measure | Analyse separately from the full ANC composite |
The broader principle is direct. An indicator name is not a metadata record. Before charting a trend, an analyst needs the questionnaire wording, numerator, denominator, event window, and tabulation note for each round.
This is particularly relevant when using IIPS Mumbai reports trend analysis tables. Published reports are indispensable, but summary tables are the end of a methodological chain, not the beginning. The table title may be stable while the respondent base, recall period, or geographic coverage has shifted underneath it.
Reference periods can create false movement
Reproductive health indicators are often event-based. Births, pregnancies, antenatal visits, child deaths, contraceptive use, and immunisation are measured over defined periods. When those periods change, estimates may no longer represent equivalent exposure windows.
The available survey review material indicates that DLHS-4 collected information on all pregnancies in the preceding five to six years. Across the DLHS series, birth-history windows varied, ranging from preceding-three-year periods to more lifetime-oriented reporting structures.
This matters most for fertility and mortality analysis.
A measure derived from births in the last three years is concentrated in a recent interval. A measure derived from a five- or six-year pregnancy history incorporates earlier service conditions, earlier mortality risks, and more distant recall. If an intervention expanded rapidly near the end of the observation period, the longer-window measure will dilute the visible effect. If reporting quality deteriorates with recall length, the longer window may introduce a different bias entirely.
The same applies to maternal care. A woman reporting on a birth six years earlier is not reporting under the same recall conditions as a woman reporting on a birth in the preceding three years. A difference between rounds can therefore combine actual health-system change with recall decay.
No universal conversion factor solves this. The analyst must align the reference periods where possible or state that the estimates represent different event windows.
For mortality indicators, the standard should be stricter. Small counts, sampling variability, and differential recall can produce unstable district-level rates. A district estimate is a survey estimate, not a complete civil registration count. Its precision depends on sample size, event frequency, survey design, and the distribution of the sampled population.
District estimates require an uncertainty framework
DLHS-3 covered approximately 720,320 households in 601 districts. District household sample targets were set at 1,000, 1,200, or 1,500 households according to district performance categories.
These are substantial survey operations. They are not censuses.
A district prevalence of 42 percent and another of 46 percent may indicate a meaningful difference. It may also be compatible with sampling error, depending on the underlying sample, design effect, subgroup restriction, and indicator prevalence. The issue becomes sharper when analysts disaggregate by age, caste, rural residence, parity, wealth category, or recent-birth status. The nominal district sample rapidly contracts.
A national or state aggregate can absorb some local imprecision. A district trend cannot. Two point estimates from separate cross-sectional rounds contain uncertainty from both samples. A visual increase on a map is not evidence that the underlying district population changed unless the survey design has been considered.
The DLHS-3 data-quality assessment identified age-reporting errors, incomplete or inconsistent responses, and skipped reproductive and sexual-health questions. It specifically cautioned that age misreporting affected women of reproductive ages. This has direct implications for trends around age eligibility boundaries.
The ages 15, 44, and 49 are not merely demographic labels. They are survey inclusion thresholds. Misreporting at those boundaries can move respondents into or out of the analytic universe. The effect can be small at the population level and consequential for narrow subgroups, particularly when an indicator is calculated from a restricted set of respondents.
An apparent change among women aged 15–44 may therefore contain at least four components:
- a real change in behaviour, service coverage, or fertility exposure;
- a change in the survey’s eligible population;
- sampling variation between independent cross-sectional samples; and
- age misclassification or other non-sampling error.
The data alone do not automatically allocate the observed difference across those components. Claims should remain proportionate to that limitation.
A workable harmonisation protocol
The practical objective is not perfection. It is to prevent invalid comparisons from acquiring the authority of a time series.
A disciplined harmonisation protocol can be applied indicator by indicator.
First, write a precise target estimand. “Maternal health trend” is not an estimand. “Percentage of births in the preceding three years to currently married women aged 15–44 occurring in an institution, among non-AHS states” is. It is narrower, but it can be audited.
Second, construct a metadata ledger for each survey round. The minimum fields are respondent universe, age range, marital-status criteria, geography, numerator, denominator, event reference period, questionnaire wording, weighting approach, and known data-quality limitations.
Third, identify the common analytic intersection. If DLHS-2 covers currently married women aged 15–44 and DLHS-3 offers that subgroup, use it. If DLHS-4 excludes household data from AHS states, restrict earlier rounds to non-AHS states for that comparison. The common denominator is often smaller than the published headline universe. That is a loss of coverage, but a gain in validity.
Fourth, separate directly comparable results from directional evidence. Some measures can support a formal comparison after harmonisation. Others can only show that conditions were measured differently across time. Those categories should not be merged in the same chart without explicit annotation.
Fifth, report uncertainty. At minimum, describe district estimates as survey-based and avoid treating small changes as decisive. Where design information permits, calculate standard errors and confidence intervals using the relevant survey weights, clustering, and stratification. A point estimate without its uncertainty is a partial result.
Finally, preserve the original label. If a source reports safe delivery, write safe delivery. If it reports institutional delivery, write institutional delivery. If the analysis uses a restricted non-AHS population, state that restriction in the figure title and text. Precision in naming prevents inflation in interpretation.
The implication for population analysis
The DLHS series remains one of India’s most useful district-level sources for reproductive and child health analysis. Its value is not reduced by methodological discontinuity. Its value depends on recognising it.
The appropriate analytical model is not a single uninterrupted national trend line from 1998–99 to 2012–13. It is a set of linked survey windows with different respondent universes, geographic frames, and reference structures. Some indicators can be bridged. Some can be bridged only within matched subpopulations. Some require a break in the series.
The policy consequence is straightforward. Before attributing a change to programme performance, service expansion, or demographic transition, reconstruct the comparison population. Align the denominator. Match the geography. Verify the indicator definition. Inspect the event window. Quantify uncertainty.
Without those steps, DLHS trend analysis can produce a clean chart and an incorrect conclusion. With them, the series can still identify demographic shifts and service gaps at a scale that administrative data rarely match.