DLHS district boundary shifts: avoiding mapping traps
India moved from 339 districts in 1961 to 593 in 2001, 640 in the 2011 Census, and roughly 720 by 2020.

Any longitudinal district series that treats those administrative units as stable observations is structurally compromised before the first indicator is calculated.
This is the central problem in DLHS district boundary mismatch longitudinal analysis. DLHS rounds, Census tables, and later survey systems may attach the same district label to different land areas, populations, urban shares, and service infrastructures. A map can still render cleanly. A regression can still converge. Neither result establishes comparability.
The error is usually mundane. A researcher merges files on district name. Or applies a current district shapefile to an older survey round. Or interprets a fall in fertility, antenatal care coverage, or child morbidity as a temporal change when the underlying denominator has been administratively redrawn. The result is not necessarily noise. It can be a directional bias.
A district is not a fixed unit merely because its name survives from one survey round to the next.
The anatomy of administrative volatility
District formation is an ordinary feature of Indian administration. It changes the analytical geography of population health.
The increase from 339 districts in 1961 to 593 in 2001 was not a cosmetic enlargement of a directory. It created smaller administrative units with different population compositions. The process continued: 640 districts were recorded in the 2011 Census, and estimates for 2020 place the total at approximately 720. The exact current count depends on the reference date and administrative source; delayed census operations make a single contemporary number less useful than it appears.
For a DLHS analyst, the relevant question is narrower: did the geographic support of observation d in survey round t match the geographic support of observation d in survey round t + 1?
In many cases, it did not.
Boundary changes affect at least four parts of the measurement system:
- Population denominators. A parent district loses villages, towns, or blocks. Its total population and age structure change even if every household-level rate within the retained territory is unchanged.
- Urban-rural composition. Redistricting may transfer rural or peri-urban areas into a previously urban-dominated district, shifting aggregate reproductive health indicators through compositional change alone.
- Administrative service geography. District hospitals, health facilities, programme management units, and reporting chains may be reorganised along with borders. A change in reported service coverage can therefore combine real programme performance with a changed catchment population.
- Sampling and weighting context. Survey estimates are built from sampled households within a specified administrative geography. Reassigning those observations after the fact without an explicit geographic crosswalk does not recreate the original sampling frame.
The scale of the mismatch is measurable. In comparisons between DLHS-3, fielded in 2007–08, and NFHS-4, fielded in 2015–16, only 506 of the 640 districts listed in the 2011 Census had common geographical boundaries suitable for direct comparison without harmonisation. That is not a marginal data-cleaning issue. It removes more than one-fifth of nominal district observations from a strict like-for-like panel.
A district-level dashboard that reports all rows as continuous series should therefore disclose its geographic rule. “District trend” is not a method. It is a label.
The parent-child fallacy
The most persistent error is the assumption that a retained district name identifies a retained district geography.
A parent district is split. The parent name remains in the administrative record. The new child district receives a new name. A basic merge now finds the parent district in both rounds and reports a successful match. This is false continuity.
The post-split parent is a residual unit. It is not the pre-split parent minus a negligible adjustment. It may have lost a concentrated tribal population, a growing town corridor, a remote block with poor facility access, or a cluster of villages with high fertility. Its observed change across rounds is a mixture of temporal change and territorial subtraction.
This matters acutely for indicators with strong geographic gradients:
- total fertility rate proxies and parity distributions;
- modern contraceptive prevalence;
- institutional delivery and skilled attendance;
- full antenatal care coverage;
- child immunisation;
- stunting, wasting, and underweight prevalence;
- neonatal, infant, and under-five mortality estimates where sample precision permits district interpretation.
Suppose a district’s rural blocks are transferred to a new district between survey rounds. The remaining parent district becomes more urban in composition. Its institutional delivery rate may rise, and its average family size may fall, without an equivalent shift in household behaviour or clinical access. The trend is real in the revised district. It is not automatically real for the original population.
Delhi’s 2012 boundary changes illustrate the mechanism. Rural components were added to Central Delhi and New Delhi, both previously classified as purely urban in the 2011 Census framework. Any comparison that treats the labels “Central Delhi” or “New Delhi” as stable urban units across the boundary change can manufacture shifts in urban-rural composition. The error then propagates into every indicator correlated with residence.
The practical consequence is blunt. District names should never be used as the sole linkage variable in longitudinal demographic analysis in India. Names are display fields. They are not geographic identifiers.
| Apparent match | What may actually have changed | Analytical consequence |
|---|---|---|
| Same district name in two survey rounds | Parent district was split and retained its name | False trend caused by population loss |
| New child district appears after a split | Territory removed from one parent district | No direct baseline estimate for the child unit |
| Same name, altered state association or code | Administrative recoding or spelling variation | Failed or incorrect merge |
| Same district label in different states | Duplicate place name | Cross-state contamination in automated joins |
| Current shapefile applied to historic data | Present boundary imposed on past estimates | Map looks valid but geography is anachronistic |
The correct question is not whether labels match. It is whether the underlying villages, census units, or sampled clusters can be mapped to a common geographic definition.
Not every split has one parent
A second mapping trap is treating all district creation as simple bifurcation. The assumption is convenient because it permits a crude rule: assign the child district back to its parent, then aggregate. The rule fails when the child is assembled from multiple parent districts.
Udalguri in Assam is a useful case. Created in 2003, it incorporated 808 villages from Darrang and 19 villages from Sonitpur. Its historical geography cannot be recovered by assigning it wholly to Darrang. Such an assignment imports territory from Sonitpur into the wrong baseline unit. If the transferred villages differ in ethnic composition, rurality, health infrastructure, or fertility patterns, the resulting error is substantive rather than clerical.
Multi-parent formations create three distinct problems.
The missing allocation problem
To reconstruct a historic equivalent of the new district, the analyst needs the allocation of constituent lower-level units—villages, tehsils, blocks, or other officially defined components—to each parent. A district-level total from the pre-creation period is insufficient. It contains people and outcomes from territory that did not become part of the child district.
The denominator problem
Population-weighted aggregation is often used when combining district estimates. It is defensible only when the weights correspond to the population of the transferred territory, not the whole parent district. Using parent-level population totals for a partial transfer is an approximation with an unknown error term.
This distinction is particularly material for rates. A district estimate of institutional delivery cannot be split accurately by assigning a proportion of the parent rate unless the allocated component has comparable delivery behaviour and an appropriate denominator. That condition should not be presumed.
The survey design problem
DLHS estimates derive from survey samples, not from administrative registers with universal village-level coverage. Even if a boundary notification identifies transferred villages, a published district-level survey estimate may not provide enough micro-geographic detail to reconstruct a reliable historic value for the new district. The appropriate response is sometimes to suppress the comparison, not to impute precision.
Where lower-level allocation data are absent, a reconstructed district trend is an estimate of geography, not only an estimate of health.
This is where comparing DLHS rounds district data becomes a methodological exercise rather than a spreadsheet task. The analyst must distinguish between observations that are directly comparable, observations that can be harmonised by aggregation, and observations that should remain non-comparable.
Build the panel before estimating the trend
The defensible unit of analysis is often not the reported district. It is a harmonised region with constant boundaries across the study period.
The standard approach is amalgamation. Parent and child districts affected by reorganisation are grouped into composite regions that preserve the same outer boundary in every round. Rather than forcing an old district estimate into a new district frame, the analyst combines the relevant units until the geographic footprint is stable.
Historical work has used this strategy to construct 232 consistent regions for a 1961–2001 district panel. The number itself is not a universal template. It demonstrates the principle: the count of usable longitudinal units may be substantially lower than the count of named districts in any single survey round.
A workable protocol has five stages.
1. Fix the target boundary vintage. Decide whether the panel will use a baseline geography, an endline geography, or composite constant-boundary regions. This decision should precede indicator extraction. It cannot be repaired cleanly after modelling.
2. Create a district master table. Each record requires state, district name as published in each source, source-specific district code where available, survey round, boundary status, parent district or districts, and a harmonised region identifier. Preserve original labels. Do not overwrite them with standardised names.
3. Classify every linkage. Use categories such as direct match, one-to-one renamed unit, parent-child split, multi-parent formation, merged district, uncertain linkage, and duplicate-name risk. A binary “matched/unmatched” field conceals too much.
4. Aggregate only on a defined basis. Counts can often be summed when component geography is known. Rates require denominator-aware aggregation. For a proportion \(p_i\) across component districts, the composite estimate should be based on relevant denominators, not an unweighted mean of district percentages. If denominators are unavailable, the limitation belongs in the output table.
5. Retain a comparability flag in the final analytical file. Every estimate should carry a field indicating whether it is direct, harmonised, partially reconstructed, or non-comparable. That flag should travel into maps, models, and published appendices.
The aggregation rule is simple in principle but routinely mishandled in practice. If two components have 60% and 80% institutional delivery coverage, the composite is not necessarily 70%. It is 70% only if the relevant delivery denominators are equal. If one component has four times as many births, its rate should dominate the combined estimate.
For survey-derived indicators, there is a further constraint. Published district tables may not provide the numerator, denominator, weighted count, and design variance needed for exact recomputation. In that case, combining point estimates can provide an exploratory descriptive value, but it should not be represented as an original survey estimate with a valid confidence interval.
Direct comparison and harmonisation are different operations
| Method | Appropriate use | Principal strength | Principal limitation |
|---|---|---|---|
| Direct district match | Boundaries demonstrably unchanged | Preserves published district estimates | Excludes reorganised districts |
| Parent-child amalgamation | Split districts can be grouped into a stable composite | Produces consistent geographic support | Reduces spatial resolution |
| Lower-level reconstruction | Village, block, or cluster allocation is available | Can approximate a chosen boundary vintage | Data-intensive; may conflict with survey design |
| Crosswalk-based allocation | Administrative concordance provides proportional mapping | Useful for descriptive series | Allocation assumptions may dominate results |
| Exclusion with documentation | Geography cannot be reconciled | Avoids fabricated precision | Reduces coverage and may affect representativeness |
The choice is not between a complete panel and an incomplete panel. It is between an explicitly bounded panel and an apparently complete panel with unknown geographic error.
Codes, names, and the false security of shapefiles
District code changes are often treated as a technical nuisance. They are more consequential than spelling differences but less informative than researchers assume.
A source-specific code can identify a unit within a particular survey or administrative file. It does not, by itself, prove continuity across time. The code may be revised, reassigned, omitted, or attached to a district whose area has changed. Work involving IIPS Mumbai district code changes should therefore retain the code as evidence, not elevate it into a geographic truth.
The master linkage should use a compound key at minimum:
- state or union territory identifier;
- district name as reported in the source;
- survey round or census vintage;
- source-specific district code;
- harmonised geographic identifier;
- boundary-comparability status.
State identifiers are indispensable because district names repeat. At least six duplicate district names create routine risks in automated joins: Aurangabad in Bihar and Maharashtra; Bilaspur in Chhattisgarh and Himachal Pradesh; Bijapur in Chhattisgarh and Karnataka; Hamirpur in Himachal Pradesh and Uttar Pradesh; Pratapgarh in Rajasthan and Uttar Pradesh; and Balrampur in Chhattisgarh and Uttar Pradesh.
A name-only join can therefore produce a complete-looking table with records assigned to the wrong state. This error may survive basic quality checks because the joined district exists, the indicator values are plausible, and the map renders without a warning.
The same caution applies to India district boundary shapefiles used in DHS or DLHS-related work. A shapefile has a publication date, a source, a projection, an attribute structure, and an implicit boundary vintage. It is not a neutral container.
Before attaching survey estimates to polygons, establish the following:
- Which administrative reference date does the shapefile represent?
- Does it include districts created after the survey fieldwork?
- Are district names and state codes aligned with the survey’s published geography?
- Does a polygon represent a post-split parent district or the pre-split parent district?
- Are disputed, missing, or merged areas handled consistently across all rounds?
- Has the file been checked against administrative notifications or a documented concordance table?
A visually sophisticated choropleth can obscure these failures. More detailed polygons do not create more detailed historical data. They can instead impose present-day administrative logic on past household observations.
For longitudinal mapping, stable composite polygons are often preferable to a sequence of changing district maps. They sacrifice local granularity. They preserve interpretability. That is the appropriate trade.
What to report with a harmonised DLHS series
A credible output should make its geographic assumptions visible without turning the main result into a metadata dump. The minimum reporting standard is compact.
State the survey rounds, the boundary reference adopted, the number of direct matches, the number of harmonised composite regions, and the number of excluded or uncertain units. Identify whether rates were recomputed from weighted numerators and denominators or aggregated from published estimates. If confidence intervals are reported, specify whether they reflect the original survey design after harmonisation or only the uncertainty published for component estimates.
This is especially relevant when comparing DLHS-3 with later data systems. DLHS-3 was fielded in 2007–08. Later comparisons may draw on NFHS-4 in 2015–16 or other sources with different sampling frames, questionnaire designs, indicator definitions, and reporting conventions. Boundary harmonisation solves one source of non-comparability. It does not solve all of them.
The analytical sequence should therefore be ordered correctly:
1. establish indicator equivalence;
2. establish geographic equivalence;
3. assess survey-design and precision differences;
4. estimate change;
5. interpret the observed pattern.
Reversing the order produces attractive but unreliable district rankings. A district may appear to be an outlier because it experienced a genuine demographic transition. It may also appear to be one because its territory was redrawn, its urban-rural mix changed, or its published estimate was joined to an incompatible polygon.
The distinction is not academic. District estimates are routinely used to allocate attention, identify lagging reproductive health indicators, and frame sub-state policy targets. Geographic misclassification can redirect that attention toward the wrong population.
The usable map is smaller than the available map
India’s district system is dynamic. Longitudinal population analysis must acknowledge that dynamism as part of the data-generating process.
The temptation is to retain every district in every round and repair gaps with names, codes, and current maps. That produces maximum apparent coverage and minimum geographic discipline. The stronger design accepts a smaller comparable panel where necessary, constructs amalgamated regions where possible, and marks irreducible uncertainty where neither option is valid.
For DLHS district boundary mismatch longitudinal analysis, the policy implication is direct. District trend systems should publish a boundary concordance alongside their indicator tables and distinguish direct estimates from harmonised estimates. Without that separation, sub-state change is not measured. It is inferred from shifting administrative containers.
The future of district-level reproductive and child health analytics is not a larger map. It is a more stable denominator.