Population Analytics: how to compare your options
India’s demographic data systems do not measure the same population, at the same geographic scale, or on the same timetable. The Census enumerates the entire population once a decade.

The Sample Registration System monitors vital events continuously through a sample of approximately 8.7 million people. NFHS and DLHS generate household and health estimates through survey designs whose district coverage, respondent groups, and biomarker protocols differ by round.
Treating these sources as interchangeable produces false precision. A district fertility estimate from a household survey is not equivalent to an annual birth-rate estimate from the SRS. A Census population count is not a prevalence survey. NFHS-5 registration data cannot be read as a direct administrative validation of the Civil Registration System. The first step in population analytics is therefore not selecting the largest dataset. It is identifying the estimand, the population denominator, and the geographic unit that the analysis must support.
The practical question—how to compare your options—has a strict statistical answer: compare the data-generating processes before comparing the numbers.
Enumeration versus estimation: what the Census and surveys actually measure
The Census of India is a complete enumeration. Its principal strength is geographic coverage. It provides the most granular subnational population data because it attempts to count the entire population rather than infer population characteristics from a sample. For district planning, population denominators, household distributions, and broad demographic structure, that distinction is decisive.
The limitation is temporal frequency. A decennial census is not a continuous monitoring instrument. Population change between enumeration points must be estimated or supplemented with other systems. Fertility, mortality, migration, and household composition can shift materially during the interval. A census table can establish the size and structure of a population at a defined reference point; it cannot by itself provide an annual series of current reproductive or mortality outcomes.
Survey systems solve the frequency problem by accepting sampling error. NFHS, DLHS, AHS, and related programs observe selected households and use survey weights to produce population estimates. Those estimates are subject to confidence intervals, design effects, nonresponse, questionnaire effects, and the quality of the sampling frame. A point estimate without its sampling context is an incomplete result.
The distinction can be summarized as follows:
| Dimension | Census of India | NFHS | DLHS / AHS | SRS |
|---|---|---|---|---|
| Basic design | Complete enumeration | Large-scale household sample survey | District-oriented household survey; AHS included additional health measurement components | Continuous sample registration with dual-record methods |
| Primary strength | Geographic completeness and population denominators | Reproductive, maternal, child-health and household indicators | District-level health and service-use analysis, with clinical and biomarker components in relevant rounds | Annual fertility and mortality statistics |
| Geographic resolution | Highly granular subnational coverage | District estimates available from NFHS-4 onward | Designed for district-level analysis in relevant rounds | Sample-based estimates, with precision varying by geography |
| Temporal pattern | Decennial | Multi-round, not annual | Periodic survey rounds | Continuous system with annual reporting |
| Main analytical risk | Age and population structure become outdated between rounds | Sampling and comparability across rounds | Differences in instruments, cohorts, and measurement protocols | Sampling limitations and reporting lag |
| Typical use | Denominators, enumeration, structure | Health and demographic prevalence estimates | District health-system and household indicators | Vital-rate monitoring |
The choice depends on the question. If the denominator is wrong, a sophisticated model will only produce a more elaborate error. If the outcome is district-level anaemia, antenatal care, child immunisation, or contraceptive use, the Census is a poor substitute for a health survey. If the outcome is annual mortality, a periodic household survey is not a replacement for a continuous registration framework.
A dataset is not “better” in the abstract. It is better or worse relative to the estimand, denominator, geography, and reference period.
Denominators determine interpretation
Population analytics often fails at the denominator stage. A fertility rate requires a defined population at risk and a defined time interval. A maternal-health coverage indicator requires a specified subgroup, usually tied to recent births or pregnancies. A mortality indicator requires an event definition, an exposure period, and a population base.
The same district can generate different-looking rates because the sources use different denominators:
- A Census-based ratio may use the enumerated population at a particular census reference date.
- NFHS may report an indicator for women aged 15–49, children under five, households, or recent births.
- DLHS-4 and AHS included clinical, anthropometric, and biochemical measurements among adults aged 18 or older.
- NFHS-4 focused primarily on women aged 15–49 and included a 15% subsample of men aged 15–54 for selected components.
- SRS fertility and mortality statistics arise from registered events in a sample population rather than from a household prevalence questionnaire.
These are not minor technical differences. They define the population to which the result can be generalized.
SRS: the continuous system for fertility and mortality
The Sample Registration System occupies a different position from the household surveys. It is the primary continuous source of annual vital statistics in India and uses a dual-record system. The system covers an approximate sample population of 8.7 million people. Its purpose is not to measure every dimension of reproductive and child health. Its purpose is to track births, deaths, fertility, and mortality through continuing registration and verification procedures.
This makes SRS particularly useful when the analytical requirement is temporal continuity. Researchers examining long-run changes in the crude birth rate, death rate, infant mortality, or total fertility rate need a source that can support repeated annual comparisons. A household survey conducted once every several years cannot provide the same temporal structure.
SRS also has limitations that become more visible when estimates are disaggregated. Sample size and event counts affect the precision of state and smaller-area estimates. Rare outcomes can produce wide confidence intervals. Maternal mortality estimates are especially sensitive to the number of observed events and the reliability of cause-of-death classification. The exact degree of under-reporting in smaller states should not be assumed without examining the relevant SRS methodology and uncertainty measures.
The system’s dual-record design is analytically valuable because it does not rely on a single channel of reporting. Events identified through separate mechanisms can be compared and reconciled. That does not turn the resulting estimates into a complete administrative census of every vital event. It creates a continuous sample-based estimation system.
SRS against administrative registration
Administrative registration data and survey-based estimates frequently diverge. That divergence should be investigated, not automatically resolved in favor of one source.
One documented example concerns death registration. CRS-SRS-based assessments suggested that death registration in India reached 100% in 2020, while NFHS-5 data indicated a figure closer to 70%. These values cannot be placed in a single ranking without examining the question wording, reference period, sample design, record verification, and definitions of a registered death.
Several mechanisms can produce the discrepancy:
1. Different units of observation. An administrative assessment may use registered events and official reporting systems. NFHS records household responses about registration status.
2. Different reference periods. The survey may ask about deaths or registration events within a defined period, while administrative totals may correspond to calendar-year reporting.
3. Different coverage. The CRS can contain records from a broad administrative system, whereas NFHS estimates are derived from sampled households.
4. Recall and documentation effects. Respondents may not possess certificates or may interpret registration status differently.
5. Survey weighting and sampling variance. A national estimate from NFHS remains an estimate with a confidence interval, not a direct count of all registered deaths.
A discrepancy is therefore a signal about measurement systems. It is not, by itself, proof that one system is accurate and the other is defective.
District-level granularity: the key break between survey rounds
District-level analysis is often the stated objective and the hidden constraint. Not every national survey produces reliable district estimates. The ability to identify a district in the dataset is not the same as having a statistically defensible district-level estimate.
The Census offers the broadest geographic enumeration. DLHS was designed around district-level health and household measurement. The Annual Health Survey also supported district-oriented analysis in its relevant coverage. NFHS acquired district-level estimation at scale beginning with NFHS-4. Earlier NFHS rounds should not be treated as if they all provide the same district-level resolution.
The survey timeline matters:
- NFHS-1: 1992–1993, establishing an early national baseline.
- DLHS-3: 2007–2008, providing district-oriented health and household data.
- DLHS-4 and AHS: 2012–2014, with clinical, anthropometric, and biochemical measurements among adults aged 18 or older in the relevant components.
- NFHS-4: 2015–2016, marking the large-scale introduction of district-level estimates within the NFHS series.
- NFHS-5: 2019–2021, the most recently completed NFHS round in the supplied evidence base.
A comparison across these rounds is possible, but not automatically valid. The analyst must inspect whether the indicator retained the same definition, whether the respondent age range changed, whether the sample design was comparable, and whether the district boundaries remained stable.
District estimates require more than a district label
A district-level table can conceal unstable estimates. Typical household sample sizes in DLHS-3 and NFHS-4 were approximately 1,000 to 1,800 households per district. That scale may support many common indicators, but it does not guarantee narrow confidence intervals for rare outcomes or small subgroups.
Precision depends on more than the nominal number of households. It also depends on:
- The number of sampled clusters.
- The distribution of households across urban and rural strata.
- The intracluster correlation of the outcome.
- The design effect.
- The number of respondents contributing to the indicator.
- The extent and pattern of nonresponse.
- The prevalence of the outcome.
- The number of children, births, pregnancies, or deaths represented in the denominator.
An indicator reported for all surveyed households may have a substantially larger effective sample than an indicator restricted to recent births or a biomarker subsample. A district estimate of contraceptive use among eligible women and an estimate of a rare maternal outcome do not carry the same statistical stability merely because they appear in the same district file.
For serious comparative work, retain the confidence interval, standard error, sample count, and weighting information alongside the point estimate. Suppressing those fields converts a survey result into a misleading rank.
District-level availability is a property of the sampling design, not a decorative column in a spreadsheet.
Comparing NFHS and DLHS without collapsing their differences
NFHS and DLHS are often grouped together because both support reproductive, maternal, and child-health analysis. Their overlap is real. Their equivalence is not.
NFHS is a multi-round, large-scale survey coordinated by the International Institute for Population Sciences in Mumbai. Its program has developed into a central source for fertility, family planning, maternal care, child health, nutrition, household conditions, and selected biomarkers. NFHS-4 and NFHS-5 are particularly valuable for district-level comparisons, provided that the indicator definitions and sample designs are aligned.
DLHS was more directly oriented toward district-level health-system and household measurement. Its instruments and respondent structures differ from NFHS. DLHS-4 and AHS also included clinical, anthropometric, and biochemical data among adults aged 18 or older. NFHS-4, by contrast, focused primarily on women aged 15–49 and a 15% subsample of men aged 15–54 for selected measures.
This difference affects the interpretation of morbidity indices, nutritional measures, and biomarker prevalence. A haemoglobin estimate derived from one age structure cannot be compared with an estimate from another without age standardisation or a defensible explanation of the population shift. The same applies to blood pressure, anthropometry, chronic disease indicators, and reproductive histories.
A practical comparison protocol
When deciding whether NFHS and DLHS can be combined, compare the following variables in sequence:
1. Target population. Identify the eligible age groups, sex categories, household definitions, and event histories.
2. Indicator construction. Confirm the numerator, denominator, exclusions, and recall period. “Antenatal care coverage” is not a single invariant measure if the number of visits or timing criteria changes.
3. Measurement mode. Separate self-reported outcomes from direct clinical, anthropometric, or biochemical observations.
4. Sampling design. Record the stratification, cluster structure, household selection, and weighting procedure.
5. Geographic boundaries. Reconcile district reorganisations before calculating trends.
6. Reference period. Align the survey fieldwork dates rather than relying only on publication years.
7. Uncertainty. Compare confidence intervals and effective sample sizes, not only point estimates.
8. Missingness. Examine whether nonresponse is concentrated in particular districts, demographic groups, or outcome categories.
This protocol does not eliminate all comparability problems. It makes them visible. In population analytics, visible limitations are usable limitations. Hidden ones are not.
Administrative data and survey data are different instruments
Administrative systems record service contacts, legal events, certificates, facility activity, and routine reporting. Household surveys collect information from sampled respondents and may include self-reported histories, direct measurements, and record verification. The two systems observe different pathways through the health system.
For maternal and child health, administrative data may overrepresent people who reach facilities and whose encounters are recorded within reporting systems. Household surveys can capture people outside formal service channels, but they introduce recall error, sampling error, and interviewer effects. Neither source is automatically the “real” dataset.
The analytical task is to determine which bias is relevant to the question.
If the question is facility workload, routine administrative records may be the primary source. If the question is population coverage, a household survey may provide a broader denominator. If the question is annual mortality, SRS offers a continuous sample-based framework. If the question is the size and age structure of the population, Census data remain foundational.
Why apparent contradictions can be informative
Suppose an administrative system reports high service coverage while NFHS reports lower household-reported coverage. The difference may reflect:
- Services recorded against an incomplete population denominator.
- Duplicate reporting or aggregation rules.
- Facilities serving populations outside their administrative catchment.
- Respondents using services that were not captured in the relevant reporting channel.
- Misclassification of service type.
- Differences in the survey’s reference period.
- Nonresponse concentrated among underserved households.
The correct response is not to average the two estimates. It is to map their observation processes and identify the stage at which they diverge.
A similar logic applies to birth and death registration. A registration database can count recorded events. A household survey can estimate the share of respondents reporting that an event was registered. They are related indicators, but they do not have identical error structures.
Timeliness, release lags, and the cost of waiting
Data selection is also a timing decision. A theoretically superior dataset may be unavailable when a policy or resource-allocation decision must be made.
NFHS public release has historically involved a lag of approximately 9 to 22 months from completion of data collection to release of individual-level data. SRS annual reports typically carry a lag of approximately one to two years. Census data provide unmatched enumeration detail but operate on a decennial cycle.
The result is a predictable trade-off:
- Census: highest enumeration coverage, lowest frequency.
- SRS: strongest continuity for annual vital statistics, sample-based geographic precision.
- NFHS: broad reproductive and health indicators, periodic release, district estimates from NFHS-4 onward.
- DLHS/AHS: district-oriented health measurement, with round-specific instruments and coverage.
- Administrative systems: potentially more current operational data, but dependent on reporting completeness and denominator quality.
The publication date should not be confused with the observation date. NFHS-5 covers 2019–2021; the release of findings later does not make the observations current to the release year. A policy dashboard that labels all values by publication year can create a false chronology.
The right source for the analytical task
A source-selection matrix is more useful than a universal ranking:
| Research question | Preferred starting point | Why | Main qualification |
|---|---|---|---|
| District population denominator | Census | Complete enumeration and granular geography | May be outdated between census rounds |
| Annual fertility trend | SRS | Continuous sample registration and annual vital statistics | Sample-based precision varies by geography |
| Annual mortality trend | SRS | Designed for ongoing birth and death measurement | Reporting lag and small-area uncertainty |
| District maternal-health coverage | NFHS-4 or NFHS-5, where applicable | District-level household estimates | Survey weights, recall, and indicator definitions |
| Earlier district health comparison | DLHS or AHS | District-oriented design and health modules | Instruments and target populations differ from NFHS |
| Household-level reproductive and child-health indicators | NFHS | Broad indicator portfolio and repeated rounds | Not every round supports district-level inference |
| Current service delivery operations | Administrative records | Operational frequency and facility reporting | Coverage and reporting completeness require validation |
This matrix is a starting route, not a substitute for protocol review. A study requiring district-level estimates for a small subgroup may need to pool rounds, use model-based estimation, or reduce geographic resolution. Each option changes the inferential target.
Designing a defensible comparison across time
Longitudinal comparisons are attractive because they appear to turn separate survey rounds into a trend. The arithmetic is simple. The validity is not.
A change between two estimates may reflect a real demographic shift. It may also reflect a change in questionnaire wording, field procedures, sample composition, district boundaries, biomarker protocols, or response patterns. The longer the interval, the greater the opportunity for multiple components of the measurement system to change.
A defensible trend analysis should preserve the following sequence:
1. Define the estimand before selecting the files. State whether the target is a prevalence, rate, ratio, mean, count, or coverage proportion.
2. Fix the geographic frame. Use consistent district boundaries or document a crosswalk for reorganised districts.
3. Align the population base. Do not compare women aged 15–49 with adults aged 18 or older as if they represented the same cohort.
4. Reconstruct the indicator. Use the published numerator and denominator rules for each round.
5. Apply survey weights correctly. Unweighted household counts do not represent the population unless the design permits that interpretation.
6. Estimate uncertainty. Account for clustering and stratification where the survey documentation requires it.
7. Test sensitivity. Recalculate results under alternative denominator or boundary treatments when those choices could materially change the conclusion.
8. Separate descriptive change from causal explanation. A rising or falling indicator is evidence of temporal difference, not proof of which intervention caused it.
The distinction between association and causation is not a philosophical precaution. It is a data constraint. NFHS and DLHS can identify distributional differences and temporal patterns. They do not, without an appropriate design, establish that a particular program produced the observed change.
Small areas and rare outcomes
District-level reporting can encourage overinterpretation. A district ranking may show a sharp difference in a rare outcome even when the confidence intervals overlap broadly. Conversely, a modest point difference may be statistically precise if the denominator is large and the design effect is controlled.
For rare events, use the event count and interval estimate as primary evidence. Avoid ranking districts by a point estimate alone. Where the numerator is small, consider aggregating years, combining adjacent areas only when substantively defensible, or applying a model designed for small-area estimation. Do not manufacture stability by displaying two decimal places.
The same rule applies to stratified cohorts. If a survey estimate is further divided by caste, residence, wealth quintile, age, or sex, the effective sample size declines. A national estimate can be stable while a district-by-subgroup estimate is not. The dashboard may still display both with equal visual prominence. The confidence intervals will not.
A route through India’s demographic data landscape
A sound analysis normally begins with the Census for population structure and geographic denominators, then moves to SRS for annual fertility and mortality patterns, and to NFHS or DLHS for household-level health and reproductive indicators. Administrative records add operational detail. They should be triangulated, not substituted indiscriminately.
The route is clear when the question is explicit:
- Need a population count or age-sex structure? Start with Census.
- Need annual vital statistics? Start with SRS.
- Need district-level maternal, child, nutrition, or reproductive indicators from recent rounds? Examine NFHS-4 and NFHS-5.
- Need earlier district-oriented health-service evidence or clinical measurement? Examine DLHS and AHS.
- Need current facility operations? Use administrative data, then test completeness against independent sources.
- Need a time trend? Build a harmonisation table before calculating the trend.
The primary keyword in this decision problem is awkwardly phrased, but the underlying task is straightforward: how to check how to compare your options without confusing enumeration, estimation, registration, and reporting. The answer is methodological rather than cosmetic. Every published number needs a defined population, geography, period, instrument, and uncertainty structure.
India’s reproductive and child-health evidence base is extensive. Its instruments are not interchangeable. Census, NFHS, DLHS, AHS, SRS, and administrative registration systems each illuminate a different segment of the population-health dashboard. The analyst’s responsibility is to preserve those distinctions through the final table, model, and policy claim.
The policy implication is direct. District planning should not rely on a single headline indicator when the underlying sources operate at different frequencies and geographic resolutions. Use the Census to anchor denominators, SRS to monitor vital-rate movement, and household surveys to estimate coverage and health conditions in defined cohorts. Where the sources disagree, treat the disagreement as a measurement problem to be diagnosed. That is not an inconvenience in the data. It is the data.