rchindia

Evidence-based maternal health insights across India

Demographic data paths: which Indian survey fits your study?

When I sit with a junior researcher who has just been handed a district health profile to interpret, the conversation almost always lands on the same anxiety: which numbers can they trust, which years compare to which, and where the gaps fall.

UpdatedJuly 30, 2026
Read time16 min read
Demographic data paths: which Indian survey fits your study?

In India, that anxiety is not a sign of inexperience. It reflects the genuine messiness of a demographic data architecture that has evolved in fits and starts: some surveys remain active, others sit in archives, and one foundational instrument—the decennial Census—has not yet been renewed after 2011.

Partner offers will appear here.

For anyone trying to understand reproductive and child health patterns below the state level—whether you are a clinician mapping antenatal care uptake, a policy analyst comparing nutrition indicators, or a public health student assembling a thesis—the source you choose will shape every conclusion you draw. District level demographic data sources in India do not form one neat series. They are overlapping instruments with different samples, geographies, denominators, and ambitions. That is inconvenient. It is also the central fact of the work.

The Evolution of District-Level Granularity: From Census to NFHS-4

For most of the post-independence era, India’s district-level demographic data effectively meant the Census. It is the only instrument designed to touch every household on a fixed cadence: exhaustive in scope, granular down to villages and towns, and indispensable for understanding population structure. But a census counts people. It does not ask them, in the detail required for reproductive and child health research, about fertility intentions, contraceptive use, child feeding, antenatal care, delivery practices, or the pathway to a health facility.

That gap is where the National Family Health Survey, or NFHS, enters.

What many newer researchers do not realise is that the early NFHS rounds were not designed as district-level tools. NFHS-1, NFHS-2, and NFHS-3 produced estimates reliable primarily at the state level, with selected urban-rural comparisons. They are invaluable for studying broad demographic transition, family planning, maternal care, and child survival over time. They are not, however, a licence to produce precise claims about every district in India.

NFHS-4 changed the practical landscape. Conducted in 2015-16, it introduced district-level estimates across 640 districts, making a nationwide district lens possible for a large set of reproductive, maternal, child health, nutrition, and household indicators. That shift matters more than it may seem. Before NFHS-4, district analysis often required researchers to work with Census variables, DLHS material, state aggregates, or selective programme data. After NFHS-4, a district factsheet became a recognisable unit of public-health discussion.

NFHS-5 carried that district orientation forward, covering 707 districts during fieldwork conducted between 2019 and 2021. It became the reference point for an enormous share of recent work on fertility, institutional delivery, immunisation, anaemia, nutrition, and women’s health. Researchers sometimes speak as if NFHS-5 created district comparability. It did not. NFHS-4 did the foundational lifting; NFHS-5 broadened and refreshed the terrain.

The distinction is not pedantic. If you are tracing change over the later 2010s, NFHS-4 is the baseline that lets you distinguish a real shift from a post-pandemic snapshot. If you are studying a district that changed administrative boundaries, the work becomes harder still: the name on a factsheet may be familiar, but the territorial unit may not map cleanly onto an earlier round.

The district-level dataset India relies on today was not always a district-level dataset. Knowing when that changed is half the methodological battle.

For long-run work, I would resist the urge to place every available source into a single spreadsheet and call the result a trend. A Census proportion, an NFHS estimate, a DLHS measure, and a health management information system count can all describe the same district while answering subtly different questions. Their agreement can be reassuring; their disagreement can be the finding. But neither should be treated as automatic proof that one series is wrong.

Decoding the NFHS-6 Factsheets: Navigating the 101-Indicator Shift

The release of NFHS-6 factsheets brought a familiar kind of excitement: new district numbers, new rankings, and a rush to compare them with earlier rounds before the methodological notes have had time to be read properly. The headline figures are striking. NFHS-6 covered 715 districts and a large household sample during 2023 and 2024. India’s total fertility rate remained at 2.0, below the conventional replacement-level benchmark of 2.1.

The more consequential change for researchers lies in the factsheets themselves. NFHS-6 presents 101 indicators, down from 131 in the NFHS-5 factsheets. That does not mean the survey programme has suddenly become less useful. It means the factsheet should no longer be treated as a complete inventory of every variable a researcher may want.

Some items that had become familiar in district comparison work are not present in the shorter NFHS-6 factsheet set, including the sex ratio at birth, access to clean cooking fuel, and selected infant and child mortality measures. For a maternal and child health researcher, that creates a practical divide. A question about anaemia, nutrition, hypertension screening, or current service access may now have a strong recent district source. A question about sex ratio at birth, fuel use, or child mortality may require a more careful combination of NFHS-5, Census material, Sample Registration System outputs, and programme data.

As a clinician, I find the loss of a visible sex-ratio-at-birth measure particularly unsettling. It has long functioned as a sentinel indicator in discussions of gender equity and access to care. But the correct response is not to quietly substitute an older figure and describe it as current. It is to state the gap. A district-level model for a recent period cannot claim direct observation where no current district factsheet measure is available.

What NFHS-6 appears to strengthen is the focus on nutrition, anaemia, and non-communicable disease risk. That aligns with the changing burden of disease: reproductive and child health work no longer sits in a sealed compartment separate from hypertension, diabetes risk, obesity, and women’s lifelong health. Still, an expanded emphasis is not the same as uncomplicated comparability. A variable can survive across rounds while its denominator, eligible population, recall period, or reporting convention changes. The first job is not to chart the percentage. It is to read what the percentage means.

RoundFieldwork periodDistricts coveredFactsheet indicator noteWhat it suits best
NFHS-42015-16640Introduced nationwide district-level estimatesPre-pandemic baseline and longer district comparisons
NFHS-52019-21707131 indicators in the factsheetsPandemic-era reproductive and child health snapshot
NFHS-62023-24715101 indicators in the factsheetsRecent work on anaemia, nutrition, and selected screening measures

The team at IIPS Mumbai has maintained enough continuity across rounds to make NFHS the strongest spine for many contemporary comparisons. But “enough continuity” is not “perfect identity.” Total fertility rate, contraceptive prevalence, institutional delivery, and full immunisation can often be compared meaningfully, provided the researcher checks definitions and notes. The footnotes are not decorative fine print. They are where comparability either survives or quietly disappears.

A useful discipline is to create a variable crosswalk before analysing anything: the exact indicator label, numerator, denominator, eligible respondents, recall period, geography, and round-specific caveat. This takes an afternoon. It can save months of explaining why a beautiful trend line rests on unlike measures.

The Legacy of DLHS and AHS: Managing Pooled Datasets for Longitudinal Studies

Here is where my clinical instincts kick in hard. I see well-intentioned graduate students download DLHS data, treat it as if it sits comfortably beside NFHS, and then encounter uncomfortable questions from a dissertation committee. Usually those questions are justified.

The District Level Household Survey, or DLHS, was conducted in four rounds between the late 1990s and the mid-2010s. It was designed as a major source for reproductive and child health monitoring, with a distinctly district-oriented logic. That makes it essential for historical fertility rate analysis datasets in India. It also makes it dangerous in careless hands, because its coverage and design changed across rounds.

DLHS-1 was conducted in 1998-99, DLHS-2 in 2002-04, DLHS-3 in 2007-08, and DLHS-4 in 2012-14. Meanwhile, the Annual Health Survey, or AHS, operated during the early 2010s across 284 districts in nine high-focus states: Assam, Bihar, Chhattisgarh, Jharkhand, Madhya Pradesh, Odisha, Rajasthan, Uttarakhand, and Uttar Pradesh.

During that period, DLHS-4 did not collect data in those AHS states. This was not an accidental hole in the map. It reflected a division of labour: AHS concentrated on states with major maternal, child mortality, and fertility burdens, while DLHS-4 covered the remaining geography. The design made operational sense. It creates a very real complication for anyone who later wants to call DLHS-4 a national dataset.

DLHS-4 alone cannot be treated as a standalone national picture. Nor can a researcher simply average a DLHS-4 district estimate and an AHS district estimate as if the two came from one identically designed sample. They did not. The questionnaires overlapped but were not identical; the sampling frames and analytic purposes differed; and AHS had a stronger mortality-monitoring orientation in its high-focus geography.

For work centred on the 2012-13 window, published research commonly uses pooled DLHS-4 and AHS material, applying the relevant weights and documenting the construction of the analytic file. “Pooled” should not mean “thrown together.” It should mean that the researcher has made an explicit methodological decision about how two complementary sources will represent a national or multi-state period.

DLHS and AHS answer different questions in different geographies in different years. Treating them as interchangeable is the most common methodological mistake I see at the district level.

If you are building a historical series across the 2000s, approach the surveys as temporal slices rather than as frames from a continuous film. DLHS-2, DLHS-3, and a pooled DLHS-4/AHS window can tell a powerful story about service uptake, fertility, child health, or women’s health. But the story needs pauses between scenes. Administrative reorganisation, questionnaire revisions, survey coverage, and changing health-system priorities all affect what an apparent trend is actually measuring.

A few practical cautions matter here:

  • Keep the geographic unit visible. District boundaries change. If a district split after an earlier survey round, a direct comparison with the newer district can become a false precision exercise unless the units are harmonised.
  • Separate estimates from counts. A survey proportion is not a programme count. A district may report a high share of institutional delivery in a survey while facility records reveal constraints in volume, referral pathways, or quality of care.
  • Use weights deliberately. Pooled files are not self-weighting simply because they sit in one folder. Document which weights were used, what population they represent, and whether the analysis is district-, state-, or household-level.
  • Do not overpromise continuity. The strongest longitudinal paper is often the one that says, clearly, “these periods are comparable within defined limits,” rather than claiming an uninterrupted annual trend that the data cannot support.

Data access also deserves its own line in the methods section. IIPS Mumbai reports data access pathways and archival arrangements are part of the practical infrastructure of this work, but the usable file is only the beginning. Metadata, questionnaires, sampling documentation, and factsheet notes should travel with the dataset from the first download to the final manuscript.

Methodological Constraints: Why the 2011 Census Remains a Static Benchmark

There is an elephant in every Indian demography seminar room: the 2021 Census was deferred, and the 2011 Census remains the last completed decennial enumeration. This matters more than it first appears.

Census data are not merely population totals. They sit beneath an enormous amount of district analysis. They anchor population structure, urban-rural composition, household characteristics, literacy, work patterns, and many of the denominators used in planning exercises. When a district health analysis describes coverage, burden, need, or potential caseload, it often depends—directly or indirectly—on a population base whose most complete enumeration is now from 2011.

That does not make Census 2011 useless. Quite the opposite. It remains the strongest structural benchmark available at district and subdistrict levels. It is especially valuable for variables that change relatively slowly: the broad rural-urban distribution, household size patterns, settlement structure, and population composition. The mistake is to treat it as a live count for current planning without acknowledging its age.

Migration is where this becomes especially visible. Internal mobility has changed the demographic shape of many districts through seasonal work, urban expansion, return migration during the pandemic period, and new industrial corridors. A household survey can capture pieces of mobility through respondent reports and migration modules. It cannot substitute for a full enumeration when the question is: how many people are now actually living in this district, in this town, in this settlement?

That limitation changes how one should read current rates. Coverage proportions from NFHS are survey estimates and can be highly useful on their own terms. But when they are converted into estimated numbers of women needing antenatal care, children requiring immunisation, or couples needing contraceptive services, the denominator assumption becomes decisive. A polished district dashboard can conceal a very old population base.

The comparison of SRS versus NFHS district data is also frequently misunderstood. The Sample Registration System provides vital-rate evidence that is indispensable for mortality and fertility discussion, particularly at national and state levels. NFHS provides much richer information on household conditions, reproductive behaviour, service use, nutrition, and selected health outcomes, including district-level estimates in its later rounds. They are complementary instruments, not competitors in a single race.

The 2011 Census is not old data to be discarded; it is the only complete structural picture we have, and every post-2011 district estimate should declare how it moves beyond it.

For now, the defensible approach is triangulation:

  • use Census 2011 for structural context and population baselines;
  • use NFHS for district-level reproductive, maternal, child health, nutrition, and household indicators;
  • use SRS for vital-rate context where its geography and publication level fit the question;
  • use DLHS and AHS for historical district analysis, with their coverage boundaries made explicit;
  • use programme data cautiously, as evidence of service delivery rather than a ready-made substitute for population prevalence.

No single source carries the whole story. The researcher’s task is not to crown a winner. It is to make the sources speak honestly to one another.

Strategic Data Selection: Matching Research Objectives to Survey Strengths

Let me bring this back to the question facing the researcher or practitioner with a district-level problem in front of them: which survey fits the study?

If the study is about fertility transitions, NFHS provides the clearest recent pathway. NFHS-6 offers the latest national and district-oriented frame where the relevant measure is available; NFHS-5 provides the immediate earlier comparator; NFHS-4 supplies the first nationwide district-level baseline. Census 2011 remains necessary for understanding the demographic structure beneath those rates. For a fertility story extending into the 2000s, DLHS and AHS become part of the record—but not as interchangeable copies of NFHS.

If the study is about maternal and child health service utilisation—antenatal care, institutional delivery, postnatal contact, contraceptive use, or immunisation—NFHS remains the cleanest single family of sources for recent district work. NFHS-4 introduced the district-level foundation, and NFHS-5 expanded the comparative record. Before assuming a direct NFHS-6 continuation, verify that the specific service indicator appears in the factsheet or available data and that its definition has not shifted.

If the study is about nutrition and anaemia, NFHS-6 deserves close attention. The recent factsheets place substantial emphasis on nutrition, anaemia, and selected non-communicable disease measures. Yet even here, a researcher should avoid treating a single district estimate as a final clinical verdict. Anaemia prevalence, for example, can illuminate scale and inequality, but it does not by itself explain diet, infection burden, supplementation adherence, testing practice, or quality of follow-up.

If the study is about gender and sex composition, the source choice narrows. Census 2011 remains the definitive full-enumeration anchor for cross-sectional sex composition. SRS can provide annual context at broader geographic levels. Where a recent district factsheet does not publish sex ratio at birth, the gap should be named, not filled with a confident-looking proxy.

If the study is historical, especially one spanning the early 2000s through the early 2010s, DLHS and AHS are unavoidable. Treat each round as an instrument shaped by its own time. Use pooled DLHS-4/AHS material only with a clear account of weighting, coverage, and harmonisation. The more ambitious the longitudinal claim, the more carefully its seams need to be shown.

The practical hierarchy is simple, even if the underlying data system is not:

1. Start with the research question, not the file that is easiest to download.

2. Choose the geographic level the survey can genuinely support.

3. Check whether the indicator exists in every round you intend to compare.

4. Read the questionnaire and factsheet footnotes before calculating differences.

5. State what the data cannot establish, especially where district boundaries, denominators, or definitions changed.

Three practices I urge on every researcher I mentor apply regardless of the dataset. First, open the factsheet footnotes before opening the spreadsheet. A two-percentage-point movement that looks like a trend may be a definitional change. Second, be precise in requests for IIPS Mumbai reports and data access: specify the survey round, district list, year, unit of analysis, and indicator codes where possible. Third, write methods sections as if the survey design matters—because it does. Treating rounds as interchangeable instruments may make the table easier to build, but it makes the conclusion weaker.

The architecture of Indian demographic data is more navigable than it first appears, but only if you start from the right door. Take the survey design seriously. Take the temporal window seriously. Take geography seriously. And never let one dataset quietly stand in for another simply because the variables have similar names.

The women, mothers, children, and households whose lives are encoded in these numbers deserve interpretations that honour the work of the field teams who gathered them. In district-level population analytics, that begins with a modest but demanding discipline: choose the right source for the right question, and let the limits of the evidence remain visible.

FAQ

Can I directly compare NFHS-5 and NFHS-6 indicators?
Not always. While many indicators remain comparable, NFHS-6 reduced the number of factsheet indicators from 131 to 101, meaning some previously tracked metrics are no longer available or may have changed in definition.
Is it appropriate to use the 2011 Census for current district-level population estimates?
The 2011 Census is the best available structural benchmark for broad characteristics, but it should not be treated as a live count for current planning due to significant changes in migration and urban growth since its release.
How should I handle district-level data when administrative boundaries have changed?
You must harmonize the territorial units before making comparisons. If boundaries have shifted, a direct comparison between an older and newer district may result in false precision unless the units are explicitly mapped to one another.
Can I treat DLHS-4 and AHS data as a single national dataset?
No. DLHS-4 and AHS were designed with different geographic focuses and sampling purposes; they should only be used together if you apply relevant weights and document the specific methodological decisions made to pool them.
What should I do if a specific indicator is missing from the latest NFHS factsheet?
You should state the data gap clearly in your research. Do not substitute it with an older figure or a proxy, as this misrepresents the current evidence base.