DLHS sample weights: when to normalize district data
A district health officer opens a DLHS-3 indicator table and sees a sharp gap in antenatal care coverage between two blocks. A staffing proposal begins to form around it.

Then a state-level analyst runs the same indicator from the same microdata with a different weight variable, and the gap largely disappears.
That is a hypothetical illustration, not a reported DLHS result. But the analytical problem is very real.
Same households. Same questionnaire. Same fieldwork. Different answer.
This is not necessarily a data-entry failure, and it is not a minor technical disagreement. It is what happens when a weight built for one inferential job is used for another. In district health planning, a dlhs sample weight adjustment district analysis can look entirely respectable while quietly answering the wrong question.
DLHS weights do not merely make a table “representative.” They encode selection, nonresponse, geography, and the level at which an estimate is meant to stand. A district percentage, a state percentage, and an estimated population total may all begin with the same respondent record. They do not automatically belong to the same weight.
The Mechanics of DLHS-3 Weighting: From Selection Probability to Response Adjustment
DLHS-3 weighting begins with the familiar core of probability sampling: the inverse of selection probability. Each sampled household receives a base weight reflecting how likely it was to enter the survey through the sample design.
That probability is not one event. It is a sequence: selection of the primary sampling unit (PSU), selection of a segment where segmentation applied, and selection of a household within the retained unit. Rural and urban sampling did not follow an identical sequence, so their base weights cannot be treated as though they emerged from one flat sample.
Response adjustment comes next. When sampled households do not respond, the households that do respond represent more than themselves. The household weight therefore incorporates an adjustment related to household response. Women’s weights follow the same broad logic but use the relevant adjustment for eligible women.
This distinction matters more than it first appears. A household-level indicator and a woman-level indicator may sit beside each other in a district profile, but they do not necessarily travel through the same weighting path.
| Unit of analysis | Typical example | Weight logic to verify |
|---|---|---|
| Household | Drinking-water access, sanitation, household assets | Household selection and household response adjustment |
| Eligible woman | Antenatal care, delivery care, contraceptive use | Household selection plus eligible-woman response adjustment |
| Child-related record | Immunisation or care indicators linked to children | The extract’s documented child or linked-respondent weighting rule |
| State aggregate | State-level estimates intended to align with controls | Whether a calibrated state weight is supplied |
The operational rule is plain: begin with the variable that matches the record under analysis. A household file does not become a women’s file because the outcome is about health. Nor does a women’s record inherit the household weight merely because both records came from the same sampled dwelling.
The field design also sets the scale of what the survey can credibly do. District samples were planned in household tiers, including targets of 1,000, 1,200, or 1,500 households and corresponding PSU targets of 22, 27, or 33 households. That is a design for district indicator estimation with usable precision. It is not an invitation to treat every weighted record as a miniature census count.
This is where many analyses go off course. The arithmetic of weighting is easy enough to run in any package. The substantive question is harder: what population does this particular weight represent, at this particular level of aggregation?
Normalization vs. Calibration: Why District-Level Estimates Require Specific Weights
Normalization and calibration are often discussed as though they were two labels for the same housekeeping step. They are not.
Normalization changes the scale of a set of weights. In the convention common to demographic health surveys, initial weights may be divided by their mean so that their sum equals the number of records in the relevant sample. For weighted means and proportions, that common rescaling does not alter the estimate. If every respondent’s weight is multiplied or divided by the same positive constant, a weighted percentage remains the same percentage.
That makes normalized weights useful for district estimates such as:
- the share of women receiving a specified service;
- the mean number of contacts or visits reported in the survey;
- the percentage of households reporting a facility or condition;
- comparisons between groups within the same district, provided the design is properly declared.
But normalization does not preserve population totals. A normalized district weight can tell an analyst that an estimated proportion is a particular percentage. It cannot, by itself, support the claim that this proportion corresponds to a specified number of women, households, or children in the district.
Calibration does something else. State-level household and women’s weights in DLHS-3 were derived from district weights using external population controls, so that state-level survey results would not depart from corresponding population information. This gives the state-calibrated weight a different task: reconciling estimates with a population benchmark at the level for which that calibration was constructed.
| Question being asked | Appropriate starting point | What must not be assumed |
|---|---|---|
| What proportion of eligible women in this district used a service? | District-oriented weight for the relevant respondent file | That its weighted sum equals the district population |
| What is the state-level total or distribution aligned with controls? | State-calibrated weight, where documented | That it remains calibrated for every district slice |
| How many women in one district have the characteristic? | A weight or estimation procedure with verified district control totals | That a normalized weight can be expanded into a population count |
| Is one district higher than another? | Comparable district estimates with design-based uncertainty | That a visible difference is statistically meaningful |
The distinction is not academic. A percentage and a total are different estimands. One can survive normalization unchanged; the other cannot.
Normalization fixes the scale of a weight. Calibration fixes the population relationship behind it. They should never be swapped by convenience.
A common mistake is to calculate district proportions with a state-calibrated weight simply because that variable appears more official or produces totals that look more intuitive. The result may still be numerically tidy. That is not the same as being a district-valid estimate.
The opposite error is just as tempting: taking a district-normalized weight, summing it, and calling the result an estimate of district population. The weighted sum is governed by the normalization convention, not by the number of people living in the district. A plausible-looking total is still an unsupported total.
For district analysis, the first question is therefore not “Which weight is largest?” or “Which one produces the expected denominator?” It is: was this weight built for the geography, unit of analysis, and estimand now on the table?
This is also the point at which the language used in reports needs discipline. “Estimated number of women” is a stronger claim than “weighted share of women in the sample domain.” If the analysis has only district-normalized weights and no verified district population controls, report the proportion clearly and resist manufacturing a count.
Navigating the Sampling Frame: PSU Segmentation and Rural-Urban Design Impacts
DLHS-3 used a stratified random design with different field structures in rural and urban areas. Rural areas used a two-stage design, while urban areas used a three-stage design, with the PSU frame drawn from the 2001 Census.
In rural areas, villages functioned as PSUs. In urban areas, the design included an additional census enumeration block layer. That extra stage is not a bureaucratic detail. It changes the route through which a household enters the sample and therefore changes the selection probability embedded in the base weight.
Large villages added a further layer of complexity. Villages with more than 300 households were segmented before household selection. Villages with 300–600 households were split into two equal segments. Larger PSUs were divided into segments of roughly 150 households, with two segments selected using probability proportional to size.
The practical implication is not that one can infer a universal variance pattern from PSU size. One cannot. The implication is that segmentation is part of the design history of a record. It may affect the relevant selection probabilities, the clustering structure, and the interpretation of identifiers in the extract.
An analyst should therefore verify, in the specific file being used:
- what constitutes the PSU identifier;
- whether segment identifiers are present and how they are coded;
- whether the file’s cluster variable refers to the original PSU, a selected segment, or another fieldwork unit;
- how strata are represented, especially when rural and urban records are combined;
- whether the supplied weight already incorporates the selection stages and segmentation adjustments.
Without that verification, it is unsafe to assume that every numeric cluster code describes the same sampling unit in the same way. It is equally unsafe to build a multilevel or survey-design model on the assumption that segmentation details can be reconstructed from PSU size alone.
| Design feature | District-indicator implication | Risk if ignored |
|---|---|---|
| Two-stage rural and three-stage urban selection | Selection probabilities differ across residence types | Treating rural and urban records as though they entered through one identical design |
| Segmentation of larger villages | Segment selection is part of the sample path | Omitting design information needed to interpret weights or clusters |
| 2001 Census sampling frame | Estimates refer to the survey frame and fieldwork context | Projecting results onto later boundary or settlement changes without a bridge |
| PSU household targets | Nominal sample size was set by design tier | Reading nominal cases as effective independent observations |
The 2001 frame deserves special care in contemporary district work. District boundaries, settlement patterns, and populations can change after a survey frame is built. A DLHS-3 estimate remains useful evidence about conditions observed in its reference period. It does not automatically map onto a later administrative geography merely because the district name appears familiar.
This matters when analysts join microdata to current administrative files. A clean merge can conceal a historical mismatch. If a district was split, merged, renamed, or materially redrawn, the analyst needs a stated crosswalk logic rather than an assumption that the modern label is equivalent to the survey domain.
Beyond Point Estimates: Incorporating Clustering and Stratification for Valid Variance
Producing a weighted percentage is the easy half of district analysis. Producing an honest interval around it is where the design begins to matter.
DLHS-3 is not a simple random sample of independent households. Respondents are grouped within sampled units, and the sample is stratified. Health, fertility, service access, and household conditions often have local patterning. People sampled from the same PSU can resemble one another in ways that people sampled across the district do not.
Weights alone do not solve this. A weighted estimate paired with a simple random-sample standard error is still a design-mismatched result.
For valid variance estimation, the survey-design declaration needs three ingredients:
1. the weight variable appropriate to the analytic record;
2. the PSU or cluster variable that corresponds to the sampling structure in the extract;
3. the stratum variable appropriate to the design and domain being analysed.
The exact implementation differs across Stata, R, SAS, SPSS, and other packages. The logic does not. The analysis software must be told that observations were selected through a clustered, stratified design rather than independently from one undifferentiated pool.
A district percentage without design-based uncertainty is a result. A district percentage with a naive standard error is a result wearing the wrong confidence interval.
The practical difficulty is often not statistical theory but file verification. Variable names, coding, scales, labels, and even the availability of design fields must be checked in the specific microdata extract and its accompanying materials. Analysts should not assume that fields used in one DLHS round, one recode file, or one copy of a dataset will appear unchanged in another.
That caution applies to segmentation as well. Design variables and segmentation details must be verified in the extract at hand before they are used in a survey-design declaration or a multilevel model. If a public-use file does not provide identifiers that can be aligned with the documented sampling design, the limitation belongs in the methods section. It should not be hidden behind a precise-looking standard error.
There is another recurring trap: importing conventions from another survey family without checking the DLHS file. Demographic health surveys in India and elsewhere may store standard recode weights without a decimal point, sometimes requiring a rescaling step before analysis. That convention is familiar, which makes it dangerous. It must not be assumed for DLHS merely because the surveys occupy adjacent methodological territory.
Before running a model, inspect the distribution and magnitude of the weight variable, read its label, and compare its behaviour with extract-specific documentation. A weight that needs rescaling for software convenience is not necessarily a weight that needs rescaling by a familiar factor.
The same restraint applies to singleton PSUs, subpopulation analysis, and domain estimation. If the question concerns one district, do not casually discard all other records and then declare a design without considering how the software handles strata and PSUs in the full sample. Survey packages differ in their treatment of subpopulations and lonely PSUs. The correct workflow depends on the extract’s structure and on the package’s documented design-based procedures.
Common Pitfalls in DLHS Data Processing and Population Total Extrapolation
Most expensive DLHS errors are not spectacular programming failures. They are quiet substitutions: a plausible variable chosen from a long file, an old convention imported into a new round, a district percentage stretched into a population count.
The following are the pressure points worth resolving before a district table leaves the desk.
1. Match the weight to the analytic unit. Household outcomes require the relevant household weight; women’s outcomes require the women’s weight. Child outcomes require close attention to the file structure and the documented weighting approach for those records. The questionnaire topic does not determine the weight. The unit of analysis does.
2. Identify the level for which the weight was constructed. A district-normalized weight and a state-calibrated weight are not competing versions of the same thing. They answer different estimation needs. Use the variable that matches the target geography and the reported estimand.
3. Do not turn normalized weights into district headcounts. A normalized weight may be entirely appropriate for a district proportion or mean. Its sum is not, by that fact alone, a district population total. Population extrapolation requires verified control totals or a documented population-based estimation method.
4. Declare the design before interpreting differences. A district estimate should carry uncertainty that reflects clustering and stratification. A narrow interval produced under an independence assumption is not evidence of precision; it is evidence that the survey design was ignored.
5. Keep the denominator visible. Antenatal care indicators, family-planning indicators, and household conditions can have different eligible populations. A weighted numerator paired with an unweighted or incorrectly filtered denominator can produce an answer that looks polished and means very little.
6. Treat missingness as part of the estimate. Item nonresponse, “don’t know” responses, and inapplicable categories are not interchangeable. The tabulation rule should make clear whether they are excluded, retained as a category, or handled under a documented indicator definition.
7. Do not confuse a block tabulation with a designed domain estimate. DLHS-3 was designed around districts and its documented sampling structure. A block-level result may be analytically interesting, but it does not inherit district-level reliability simply because block codes exist in the file. Small domains require caution, especially where the number of contributing PSUs is limited.
8. Record every transformation. If weights are normalized, rescaled, restricted to a domain, or otherwise transformed, the method should state exactly what happened and why. This is not clerical detail. It is the difference between a reproducible estimate and a number that cannot be audited.
The useful discipline is to separate three questions that are often collapsed into one: What share of the target population has a characteristic? How uncertain is that share under the sample design? How many people does that share represent? DLHS can inform all three questions, but not necessarily with one weight variable or one line of analysis.
A district estimate earns trust not because its decimal places look authoritative, but because its weight, denominator, geography, and variance method all describe the same inferential object. That is the standard worth holding when district data are expected to guide reproductive and child health decisions.