RCH indicator mapping: a spatial analysis roadmap
India's maternal and child health machinery operates across 766 districts, yet a single metric — the Coverage Gap Index — exposes a brutal clustering pattern: a Global Moran's I of 0.70, with 122…

India's maternal and child health machinery operates across 766 districts, yet a single metric — the Coverage Gap Index — exposes a brutal clustering pattern: a Global Moran's I of 0.70, with 122 districts locked in high-gap hotspots and 164 districts forming coldspots of relative adequacy. The geographic inequality is not random. It follows spatial autocorrelation patterns so consistent that they read like a topographic map of institutional failure. If you are working with RCH data and not applying spatial econometrics, you are reading the ledger but ignoring the footnotes where the real discrepancies live.
This is a roadmap for doing the opposite.
The Data Architecture: Four Rounds of District-Level Household Surveys
The Reproductive and Child Health programme launched in 1997, and with it came a demand for granular, district-level measurement that census data alone could never satisfy. The response was the District Level Household Survey, conducted in four rounds between 1998 and 2013, each expanding the scope and methodological sophistication of its predecessor.
DLHS-1 (1998–1999) established baseline coverage indicators for antenatal care, institutional delivery, and immunization. DLHS-2 (2002–2004) refined sampling frames and added morbidity tracking. DLHS-3 (2007–2008) scaled dramatically — approximately 700,000 sample households across 612 districts, with per-district sample sizes ranging from 1,000 to 1,500 households, stratified by district performance on key RCH indicators. The International Institute for Population Sciences in Mumbai served as the nodal agency for all four rounds, standardizing questionnaire design and field protocols across states.
DLHS-4 (2012–2013) introduced a structural innovation: a population-linked facility survey integrated directly into the household module. For the first time, researchers could cross-reference household-reported service utilization against verified infrastructure data from Sub-Health Centres, Primary Health Centres, Community Health Centres, and District Hospitals. This linkage closed a critical gap — prior rounds documented what households reported receiving, but could not independently verify whether the facility to serve them existed within a reasonable geographic catchment.
| Survey Round | Period | Key Methodological Addition | Sample Scale |
|---|---|---|---|
| DLHS-1 | 1998–1999 | Baseline RCH coverage indicators | National baseline |
| DLHS-2 | 2002–2004 | Refined sampling, morbidity tracking | Expanded |
| DLHS-3 | 2007–2008 | Performance-stratified district sampling | ~700,000 households, 612 districts |
| DLHS-4 | 2012–2013 | Integrated population-linked facility survey | District-level facility verification |
The progression from DLHS-1 to DLHS-4 is not merely incremental. Each round tightened the relationship between population-level demand data and supply-side facility data, making spatial analysis increasingly viable.
Spatial Autocorrelation: The Statistical Grammar of Inequality
District-level health data, mapped naively, produces choropleth visualizations that seduce the eye but mislead the analyst. Adjacent districts with similar outcomes may reflect genuine spatial dependence — shared infrastructure, cultural norms, economic corridors — or they may reflect nothing more than geographic coincidence. Distinguishing between the two requires spatial autocorrelation testing.
Global Moran's I is the workhorse metric. It measures whether the spatial distribution of a variable is clustered, dispersed, or random across a defined geography. A value of +1 indicates perfect positive spatial autocorrelation (high values near high values); −1 indicates perfect negative autocorrelation (high values near low values); 0 indicates spatial randomness.
For the Coverage Gap Index across Indian districts, the Global Moran's I of 0.70 (p < 0.01) delivers an unambiguous verdict: health coverage gaps cluster geographically at a level that cannot be attributed to chance. This is not a gentle tendency. It is a structural pattern embedded in the spatial fabric of service delivery.
A Moran's I of 0.70 does not suggest inequality. It confirms that inequality has a postal code.
Local Indicators of Spatial Association — LISA — disaggregate this global signal into district-level contributions. LISA classifies each district into one of five categories:
1. High-High (hotspot): A district with high coverage gaps surrounded by districts with similarly high gaps. These are the 122 districts where underperformance is regionally systemic, not locally anomalous.
2. Low-Low (coldspot): A district with low coverage gaps surrounded by similarly low-gap neighbours. The 164 coldspot districts represent regions where service delivery functions with relative coherence.
3. High-Low (spatial outlier): A high-gap district embedded in an otherwise low-gap region — an anomaly worth investigating for localized governance failures.
4. Low-High (spatial outlier): A low-gap district surrounded by high-gap neighbours — a potential model of what works under adverse regional conditions.
5. Not significant: Districts where the local spatial relationship does not reach statistical significance.
The practical implication is immediate. A national-level indicator of 75% immunization coverage tells you nothing about whether the remaining 25% is uniformly distributed or concentrated in a contiguous belt of 120 districts across three states. LISA tells you precisely that.
Bivariate Spatial Analysis: Linking Development to Health Coverage
Univariate spatial autocorrelation identifies clustering within a single variable. Bivariate spatial analysis extends the logic to two variables simultaneously, testing whether the spatial distribution of one variable correlates with the spatial distribution of another.
The canonical application in Indian RCH research: correlating the Social Development Index (SDI) with the Coverage Gap Index (CGI) across districts. A bivariate Moran's I of 0.42 (p < 0.01) indicates moderate positive spatial autocorrelation — districts with lower social development tend to cluster near other low-SDI districts, and these same districts tend to exhibit higher health coverage gaps. The correlation is real but not deterministic. SDI explains some of the spatial variance in CGI; it does not explain all of it.
This distinction matters for policy. If CGI were perfectly predicted by SDI, the intervention logic would be straightforward: invest in social development and health coverage follows. A Moran's I of 0.42 suggests a more complex causal architecture — one where health system factors (facility density, staffing ratios, supply chain reliability) operate with some independence from broader socioeconomic conditions.
The methodological workflow for bivariate spatial analysis proceeds as follows:
1. Construct spatial weights matrices. Queen-contiguity (districts sharing a boundary or vertex) or distance-band weights define the neighbourhood structure. The choice of weights matrix influences results; sensitivity testing across multiple specifications is non-negotiable.
2. Standardize both variables. Z-score transformation ensures comparability across indicators measured on different scales.
3. Compute bivariate Moran's I. The statistic measures the correlation between the standardized value of variable X in district i and the spatially lagged value of variable Y in the neighbouring districts of i.
4. Generate LISA cluster maps. Identify which districts contribute disproportionately to the global bivariate statistic — the High-High, Low-Low, and outlier categories applied to the two-variable relationship.
5. Assess statistical significance. Permutation-based p-values (typically 999 iterations) guard against false inference from spatial structure alone.
The gap between a Moran's I of 0.70 (univariate) and 0.42 (bivariate) is the space where policy operates — beyond socioeconomic determinism, within reach of targeted health system reform.
GIS Tooling and Shapefile Infrastructure
Spatial analysis of RCH indicators requires three inputs: survey microdata at the district level, district boundary shapefiles, and GIS software capable of computing spatial statistics. The first is available through IIPS data archives. The second demands careful attention to temporal consistency.
India's district boundaries are not static. Districts are carved, merged, and renamed with regularity — Andhra Pradesh's bifurcation in 2014 being the most dramatic recent example, but smaller reorganizations occur across states in most years. DLHS-3 was fielded in 612 districts; DLHS-4 covered a different district count. Any spatial analysis spanning multiple survey rounds must harmonize district boundaries to a common geography, typically the Census of India boundary definitions for the relevant reference year.
For software, the field operates without a mandated national standard. Research published in the Indian Journal of Medical Research, BMC Public Health, and related outlets has employed ArcGIS, QGIS, and GeoDa interchangeably for RCH spatial analysis. GeoDa remains the most common tool for Moran's I and LISA computation given its purpose-built spatial econometrics module and open-source licensing. QGIS handles cartographic output and geoprocessing. ArcGIS offers the most comprehensive geodatabase management for large-scale projects but at licensing cost.
| Software | Primary Use in RCH Analysis | Licensing | Spatial Statistics Module |
|---|---|---|---|
| GeoDa | Moran's I, LISA, spatial regression | Open source | Native, purpose-built |
| QGIS | Cartography, geoprocessing, shapefile management | Open source | Via PySAL plugin |
| ArcGIS | Geodatabase management, complex geoprocessing | Proprietary | Spatial Statistics Toolbox |
Shapefile acquisition for Indian district boundaries follows a standard pipeline: Census of India boundary files (available through the Survey of India or the Data.gov.in portal), pre-processed by research groups at IIPS or affiliated institutions to match survey district codes. The match between microdata district codes and shapefile attribute tables is the single most common point of failure in RCH spatial analysis. A mismatch produces silent errors — districts dropped from analysis without warning, or worse, incorrectly merged with neighbouring districts.
Childhood Stunting and Diarrhoea: Spatial Patterns in Morbidity
Beyond service coverage indicators, spatial autocorrelation has been applied to direct morbidity outcomes in the DLHS and NFHS datasets. Two indicators illustrate the approach.
Childhood stunting — defined as height-for-age below minus two standard deviations from the WHO reference median — exhibits a Global Moran's I in the range of 0.52 to 0.63 across different NFHS waves, depending on the study specification and dataset version. The variation in reported Moran's I values reflects differences in spatial weights definitions, district boundary harmonization choices, and survey wave. Regardless of the exact figure, the message is consistent: stunting clusters geographically. The high-burden districts are not randomly scattered across the Indian map. They concentrate in contiguous belts — central India, parts of eastern Uttar Pradesh, Bihar, Jharkhand, and Madhya Pradesh forming a persistent high-stunting corridor.
Childhood diarrhoea — a more temporally volatile indicator sensitive to seasonal water quality variation — shows a lower but still significant Global Moran's I of approximately 0.37. The weaker spatial autocorrelation relative to stunting reflects the different causal structure: stunting is driven by chronic nutritional deprivation and repeated infection, both of which are strongly correlated with persistent poverty and infrastructure deficits that cluster spatially. Diarrhoea incidence, while also correlated with poverty, introduces acute seasonal and environmental variation that dilutes the spatial signal.
The analytical takeaway: spatial autocorrelation strength is itself a diagnostic. High Moran's I values signal that the indicator is driven by structural, place-based factors amenable to area-level intervention. Lower values signal greater heterogeneity within regions, suggesting that intervention design must be more granular.
Methodological Constraints and the Data Gap
No discussion of RCH spatial analysis in India is complete without an honest accounting of the constraints.
Temporal lag. DLHS-4 data was collected in 2012–2013. Over a decade of demographic transition, policy intervention, and infrastructure development has occurred since. The spatial patterns identified in DLHS-4 may have shifted — some hotspots may have resolved, new coldspots may have emerged. The National Family Health Survey (NFHS) has partially filled this gap with its own district-level estimates (NFHS-4 in 2015–2016, NFHS-5 in 2019–2021), but the NFHS and DLHS use different sampling designs, questionnaire instruments, and district definitions, complicating direct comparison.
Non-spatial confounders. Spatial autocorrelation in health outcomes may partly reflect spatial autocorrelation in measurement error. Districts with weaker survey infrastructure may produce systematically biased estimates; if these districts cluster geographically (which they do, given correlations between institutional capacity and remoteness), the measured spatial autocorrelation is inflated. No published RCH spatial analysis has fully decomposed true spatial dependence from spatially structured measurement error.
Modifiable Areal Unit Problem (MAUP). Results are sensitive to the geographic unit of analysis. District-level analysis may mask sub-district heterogeneity; state-level aggregation would obscure district-level patterns entirely. The choice of district as the unit of analysis is a practical convention driven by data availability, not a theoretically grounded spatial scale for health interventions.
Facility-level spatial data. DLHS-4's facility survey provides verified infrastructure data at the point of service delivery, but the RCH Portal (Mother & Child Tracking System) captures data at the Sub-Centre level, generating unique tracking IDs for eligible couples and infants. The spatial resolution mismatch between household survey data (district-level) and facility tracking data (Sub-Centre-level) remains underutilized in published spatial analyses. Linking these datasets — mapping household-level coverage gaps against nearest-facility service delivery records — would produce a more operationally useful spatial diagnostic than either dataset alone.
The Ministry of Health and Family Welfare's Health Management Information System (HMIS) offers a third data stream, reporting routine administrative data on service delivery at facility and district levels with monthly temporal resolution. HMIS data suffers from well-documented completeness and accuracy limitations, but its temporal granularity complements the snapshot nature of survey data. Integration of HMIS time-series with DLHS/NFHS cross-sectional spatial analysis represents the logical next step for real-time RCH monitoring.
From Mapping to Intervention Design
Spatial analysis of RCH indicators is not an academic exercise. It is a targeting instrument. The identification of 122 hotspot districts for maternal and child health coverage gaps — concentrated geographically rather than scattered at random — has direct resource allocation implications. National health missions, state-level programme managers, and district collectors need to know where the gaps are contiguous and therefore where block-level or district-cluster interventions can achieve economies of scale.
The progression from DLHS-1 through DLHS-4, from univariate choropleth mapping to bivariate spatial regression, from static cross-sectional analysis to integration with HMIS real-time data, traces a methodological arc that is still incomplete. The spatial econometric toolkit exists. The data infrastructure, while imperfect, is richer than in most low- and middle-income countries. The missing component is institutional adoption — routine integration of spatial autocorrelation testing into programme monitoring, not as a research afterthought but as a standard operational dashboard metric.
India's district-level health data, when subjected to rigorous spatial analysis, reveals a landscape of clustered disadvantage that administrative averaging conceals. The 122 hotspot districts are not abstractions. They are places where children are stunted, mothers lack antenatal coverage, and immunization schedules lapse — and they sit next to each other on the map. Geography is not destiny, but in RCH programme design, ignoring it is negligence.