District Level Household Survey data variables: seven pre-analysis checks
A DLHS spreadsheet can fail before the first regression runs. Not because the analyst cannot code. Because the file is being asked to answer a question it was never designed to answer.

The recurring damage is familiar: a fertility indicator from one round gets lined up with a similarly named variable from another; household records are blended with women’s interviews; district percentages are treated as clean trends without checking who was eligible to respond. The output looks polished. The inference is broken.
DLHS is not one uniform data stream. It is a sequence of large, district-oriented survey operations with changing field periods, instruments, respondent universes, and sampling architecture. Treat it as a warehouse of interchangeable columns and the analysis will inherit every fracture in that infrastructure.
A variable name is not a definition. In DLHS, the denominator, module, round, and sample design are part of the variable.
Here is the working sequence I use before any descriptive table, map, model, or district ranking leaves the workstation.
1. Lock down the survey round and reference period
Start with the blunt question: which DLHS operation produced this record?
The field periods are not administrative trivia. They determine what population conditions the data can describe, which questionnaire logic applies, and whether a comparison is even plausible.
| Survey round | Field period | Immediate analytical implication |
|---|---|---|
| DLHS-1 | 1998–99 | Earlier eligibility and instrument structure; do not assume alignment with later rounds |
| DLHS-2 | 2002–04 | Multi-year field period; retain the round identity in every derived output |
| DLHS-3 | 2007–08 | Large district-focused operation with separate household, women’s, village, and facility modules |
| DLHS-4 | 2012–13 | Later questionnaire content includes maternal care, child care, contraception, fertility preferences, and reproductive-health knowledge |
This is the first of the seven DLHS survey data variables checklist before analysis checks because every downstream decision depends on it. A column called “institutional delivery,” “modern method,” or “knowledge of HIV/AIDS” does not become longitudinal simply because the label sounds familiar.
For every incoming file, create a round register before touching the analytical variables:
- survey round and field period;
- data file name and stated record level;
- questionnaire version and module;
- geographic coverage stated in the documentation;
- the exact denominator intended for each published indicator;
- unresolved differences between catalog metadata and technical documentation.
That register is not paperwork. It is chain-of-custody control for the analysis.
DLHS-3, for example, is generally associated with data collection from 2007–08; one dataset record specifies December 2007 to December 2008. That is historical population evidence. It cannot be casually framed as current fertility behavior, current service availability, or today’s reproductive-health coverage.
The failure mode here is fast and expensive: analysts append rounds, add a year variable, and call it a trend. What they may actually have built is a timeline of changing questionnaires and changing denominators.
2. Identify the unit of analysis before selecting variables
The next bottleneck is record level. This is where otherwise competent work starts mixing objects that do not belong in the same row.
DLHS-3 used distinct questionnaires for households, ever-married women, unmarried women, villages, and health facilities. These are not alternate windows into a single respondent. They are separate measurement systems.
The household questionnaire captures socioeconomic and household-level conditions. Women’s questionnaires carry reproductive and maternal-health information, service use, and health knowledge. Facility instruments capture the delivery side: infrastructure, staffing, drugs, equipment, and service readiness.
A clean analysis starts by naming its analytical unit in one sentence:
- Household-level question: What proportion of sampled households had a given asset, water source, or household characteristic?
- Woman-level question: What proportion of eligible interviewed women reported a specified reproductive-health or maternal-care outcome?
- Facility-level question: What share of surveyed facilities had a particular input, staffing position, drug, or service capacity?
- Geographic or district-level question: Which aggregation rule turns sampled records into a district estimate, and does the documentation support that estimate?
Do not join files merely because they share a district code. A district match is geography, not a causal or statistical bridge.
Suppose a researcher wants to test whether local facility readiness predicts antenatal care. A woman’s care record and a facility’s staffing record are not automatically a linked pair. In DLHS-3, the facility survey was population-linked: all Community Health Centres and District Hospitals in a district were covered, while selected Sub-centres and Primary Health Centres were those expected to serve the selected PSU population. That is useful design information. It is not permission to attach one facility’s stock register to one woman’s interview.
The practical fix is simple but non-negotiable:
1. Keep each module in a separate analytic table at first.
2. Build module-specific variables and denominators.
3. Document any geographic aggregation or linkage rule.
4. State whether the resulting association is individual, facility, PSU, or district-level.
5. Refuse to interpret an ecological linkage as evidence of an individual service pathway.
This is the core of district level household survey dataset preparation steps: build the data structure around the survey’s actual machinery, not around the convenience of a merged file.
3. Map the respondent universe, not just the variable label
Women’s variables are where the denominator errors become most damaging.
In DLHS-1 and DLHS-2, the women interviewed were currently married women aged 15–44. DLHS-3 expanded and changed the universe: it interviewed ever-married women aged 15–49, while also interviewing never-married women aged 15–24 through a separate instrument.
That shift changes the population represented by a measure before anyone has calculated a percentage.
A fertility preference variable among currently married women aged 15–44 is not automatically comparable with a similarly labeled item among ever-married women aged 15–49. The latter may include widowed, divorced, or separated women and five additional years of age exposure. That is not a small technical adjustment. It can move the estimate and alter the story of inequality between districts.
Build an eligibility map for every women’s indicator
For each selected variable, record these fields beside the data dictionary:
| Field to map | Why it changes the result |
|---|---|
| Questionnaire module | Separates ever-married, unmarried, household, and facility measures |
| Eligibility universe | Defines who could have answered the question |
| Age range | Alters the denominator and exposure window |
| Marital-status criterion | Can materially change fertility, contraception, and care estimates |
| Recall period | Affects comparability across rounds and outcomes |
| Skip logic | Determines whether a blank is ineligible, missing, or a true response |
| Outcome denominator | Distinguishes all eligible respondents from a narrower relevant subgroup |
This is where rch survey data variables validation becomes operational rather than ceremonial. A blank value may mean “not asked,” “not eligible,” “does not know,” “refused,” “not applicable,” or an unresolved coding convention. Those states cannot be collapsed into zero without documentation.
The inspected documentation does not supply a complete variable-level codebook, value-label list, missing-value convention, or full weighting specification for a particular downloadable microdata file. That gap matters. Before recoding, verify those details in the actual file documentation and questionnaire materials. If you cannot verify them, label the variable as provisional and do not publish a precise estimate from it.
In survey analysis, an invalid denominator does not create noise. It manufactures a false result with clean decimal places.
DLHS-4 reinforces the same warning. Its ever-married women’s questionnaire includes women’s characteristics, maternal care, immunization and child care, contraception and fertility preferences, and reproductive-health knowledge including HIV/AIDS. Those topic headings are not proof that identically named variables from DLHS-1, DLHS-2, or DLHS-3 have the same wording, coding, reference period, or universe.
4. Inspect sampling design before claiming a district pattern
District estimates are the point of DLHS. They are also where analysts most often outrun the sample design.
DLHS-3 used a two-stage stratified random sample in rural areas and a three-stage stratified sample in urban areas. The 2001 Census served as the sampling frame. Villages and urban wards were selected with probability proportional to size, and households were selected through systematic random sampling.
That architecture has consequences:
- observations are clustered, so households and respondents inside the same PSU are not independent in the way a simple random sample assumes;
- rural and urban records follow different selection pathways;
- district sample sizes were planned differently by district performance category;
- a raw sample proportion is not automatically a valid district population estimate;
- standard errors and confidence intervals require the correct weights and design variables.
For DLHS-3, planned household sample sizes were 1,500 in low-performing districts, 1,200 in medium-performing districts, and 1,000 in good-performing districts. That is operationally sensible: resources go where resolution is needed. But it also means analysts should stop treating every district result as if it carries identical precision.
A small subgroup can collapse the usable sample further. A district may have roughly the planned number of households while having too few eligible women, births, facility observations, or exposed respondents for a stable subgroup estimate. The spreadsheet will still return a percentage. The percentage does not become reliable because it exists.
Before reporting a district comparison, answer these four questions:
1. Is the estimate weighted, and has the released file’s appropriate analysis weight been verified?
2. Have PSU, strata, and any required design variables been identified from the microdata documentation?
3. What is the unweighted denominator after eligibility and missing-data rules are applied?
4. Is the result a descriptive sample figure, a design-aware population estimate, or an unsupported calculation?
If the weighting and design specification cannot be confirmed, do not claim that the figure is representative. Report the limitation plainly or stop the district comparison. There is no engineering fix for metadata that has not been recovered.
5. Trace PSU segmentation and geographic selection
PSU variables often look like back-office debris. They are not. They tell you how the field operation carved geography into measurable units.
For DLHS-3, villages with more than 300 households were segmented under specified procedures. For PSUs with 300–600 households, two equal segments were created and one was selected with probability proportional to size. For PSUs above 600 households, segments of 150 households were created and two segments were selected with probability proportional to size.
This matters for three reasons.
First, a “village” in the analytical file may represent a selected segment rather than the entire village. Treating it as complete village coverage can distort local interpretation.
Second, segmentation affects household selection probabilities. If the file includes PSU, segment, ward, or selection variables, preserve them. Do not drop them during “cleaning” because they look inconvenient.
Third, geographic IDs are not automatically stable across operations. District boundaries, district naming, coding systems, and the number of covered districts may shift between rounds or between documentation sources. A district merge is a controlled operation, not a drag-and-drop task.
The minimum geographic audit should include:
- raw district and state identifiers retained exactly as supplied;
- a separate standardized name field, never overwriting the raw code;
- an explicit crosswalk for any district split, merger, or recoding;
- a flag for records that cannot be reconciled cleanly;
- a check that PSU and segment identifiers are unique only within their proper geographic hierarchy.
Do not manufacture continuity where the sampling frame does not provide it. If the question is about change over time, a harmonized geographic unit must be defined first. Without that, a “district trend” may be tracking boundary change, coverage change, or a coding mismatch rather than population change.
6. Keep household and facility indicators in separate lanes
The DLHS-3 facility survey is unusually valuable because it puts service-delivery infrastructure next to population data within a district-focused design. It covers infrastructure, staffing, drugs, instruments, and service availability. That is exactly the material needed to identify supply chain bottlenecks and infrastructure decay.
But the linkage has to be documented with discipline.
All Community Health Centres and District Hospitals in a district were covered in the facility component. Sub-centres and Primary Health Centres were included when they were expected to serve the selected PSU population. This creates an opportunity to ask better questions than “what percentage received care?”
For example:
- Do districts with weaker facility readiness also show lower reported use of a maternal-health service?
- Is a low-coverage district constrained by household access, facility capacity, or both?
- Are stock and staffing deficits concentrated in the same geographic areas as poor care indicators?
These are valid planning questions only when the unit of inference is stated correctly. They are not proof that a specific woman failed to receive a service because a particular facility lacked a drug or staff member.
Use a layered workflow:
1. Estimate the household or women’s outcome using its own eligibility rules and survey design.
2. Summarize facility readiness using the facility sample and its documented coverage.
3. Aggregate only to the common geographic level supported by both data sources.
4. Retain separate denominators: women interviewed, households selected, facilities assessed.
5. Describe the result as a district-level association unless a documented respondent-to-facility linkage exists.
This separation prevents a common category error: confusing the demand-side record of reported care with the supply-side record of service readiness. Both are necessary. Neither can substitute for the other.
7. Reconcile coverage before publishing counts, maps, or rankings
The final check is coverage reconciliation. This is where a rushed analysis can acquire false authority from a single headline number.
For DLHS-3, documentation does not present one perfectly frictionless coverage statement. IIPS reports data collection completed in 601 districts across 34 States and Union Territories in 2008. A GHDx record describes 720,000 households in 601 districts, across all states except Nagaland. IIPS also describes an overall sample of about 700,000 households from 612 districts.
These statements may refer to different reporting universes, stages of collection, or documentation conventions. The available materials do not fully resolve the discrepancy. That is not an invitation to choose the biggest or neatest figure. It is a signal to annotate the scope of your own extract.
Before output leaves the team, reconcile:
- number of records in the supplied file;
- number of unique districts after cleaning;
- number of states and Union Territories represented;
- exclusions, including any state-level gap described in the documentation;
- difference between planned sample, completed data collection, and final analytic sample;
- records removed through eligibility rules, missingness handling, or merge failures.
Then make the scope visible in the methods note. Use language such as: “This analysis uses the available DLHS-3 extract for the documented field period and reports results for districts present after stated eligibility and data-quality filters.” It is less flashy than claiming universal coverage. It is also defensible.
A demographic health survey data cleaning checklist that skips this stage is not cleaning. It is cosmetic repair over an uninspected foundation.
The output should be an analysis-ready map, not a cleaned-looking file
The useful end product is not one enormous merged dataset with a thousand renamed columns. It is a controlled analytical package:
- a round register;
- a module inventory;
- an eligibility and denominator map;
- a geographic crosswalk;
- a sampling-design note;
- separate household, women’s, village, and facility analytic tables;
- a coverage reconciliation log;
- a list of variables that remain unverified pending codebooks or microdata documentation.
That package may feel slower than opening the file, filtering a few rows, and producing district rankings by lunch. It is slower. It is also the only route that does not turn historical survey infrastructure into misleading population claims.
DLHS data can show where reproductive and child-health systems were under strain: uneven care use, constrained facility capacity, gaps in local readiness, and district-level variation that national averages erase. But the survey will only carry that load if the analyst respects its joints. Confirm the round. Lock the unit of analysis. Build the right denominator. Trace the sample. Keep supply and demand records in their lanes. Reconcile coverage.
Then run the numbers.