Home » Measuring multidimensional poverty in countries with incomplete data

Measuring multidimensional poverty in countries with incomplete data

by NNW Bureau
0 comments

Multidimensional poverty measures offer a more complete picture of poverty than monetary measures alone, yet their adoption has been limited by fragmented data. For instance, for the most recent year, the World Bank’s Multidimensional Poverty Measure (MPM) can be estimated for only two-thirds of the world despite individual indicators covering at least 85% of the global population. This is because these indicators are scattered across different datasets, which complicates measurement and carries real policy costs, as our companion blog highlights.

Our recent working paper addresses this data challenge by combining indicators from various datasets at the national or subnational level. The main idea is simple: if one survey measures access to drinking water and another measures access to sanitation, the two can be combined to estimate the number of people deprived in both — without requiring either indicator to appear in the same dataset. This unlocks three practical applications. First, it allows multidimensional poverty to be estimated from surveys that are missing one or more indicators. Second, it makes it possible to incorporate dimensions currently excluded from the MPM due to data limitations, such as health or security. Third, it enables estimates for years in which no survey is available, including nowcasts for the present year. The rest of this blog summarizes the approach and shows that it performs well in common fragmented data scenarios where we can test it against observed data. 

Combining aggregated statistics using probability theory

Our method relies on probability theory to combine population-aggregated statistics across datasets. Unlike approaches that predict poverty at the individual or household level, it is less demanding on the data and easier to apply at a global scale. Essentially, indicators from separate datasets are combined using the chain rule under an assumption of conditional independence.

To illustrate, suppose monetary poverty (p) and access to electricity (e) are each available from two different nationally representative surveys, both of which also collect information on access to drinking water (w). If access to drinking water is conditionally independent of monetary poverty given access to electricity — that is, P(w|p,e) = P(w|e) — then, among those without electricity, the probability of lacking drinking water is the same regardless of monetary poverty status. Under this assumption, the joint probability of all three indicators can be expressed as: P(p,e,w) = P(p,e) · P(w|e), combining the joint distribution of monetary poverty and electricity (from the first dataset) with the conditional distribution of drinking water given electricity (from the second). In practice, conditional independence is a reasonable assumption, as dimensions of poverty are often interrelated and driven by common underlying factors (for example, among others see the research here and here). In a similar vein, related research shows the benefit of assuming a fixed value for the missing part of the joint distribution.

Validating the method with over 500 surveys spanning more than three decades

To validate the method, we use 571 surveys spanning 1989 to 2024 from the World Bank’s Global Monitoring Database (GMD), a collection of harmonized surveys used to report global poverty, shared prosperity, and inequality estimates in PIP, and the source for the MPM. Each survey contains information on all three dimensions of the MPM: monetary poverty, education, and access to basic infrastructure. Hence, we can estimate the “true” MPM from these surveys.

We compare the “true” MPM calculated from each survey against estimates produced in scenarios where indicators at the population group level are drawn from separate datasets. For example, the scenario p-c-r-e-w-s assumes each MPM indicator — monetary poverty (p), completed primary education (c), enrolled in school (r), and access to electricity (e), water (w), and sanitation (s) — comes from a different dataset. A more realistic scenario, pcre-ws, places monetary poverty, education, and electricity in one dataset and water and sanitation in another. We then combine the probabilities estimated from separate datasets using our method and compare it with the “true” value. The scenarios tested reflect the most common missing data patterns observed in the GMD.

Fusion performs better as more indicators overlap

Figure 1 shows the difference between the “true” MPM and estimated MPM across six data scenarios, with fusion applied at the national level and at finer population-aggregate levels: rural/urban, income quintiles, or the lowest representative level in the survey. Results for additional metrics are also shown, including an application of the Alkire-Foster method to the MPM.

As expected, the largest errors occur when no indicators overlap across datasets. Errors decrease substantially as more indicators are shared. For instance, the most common missing data pattern in the GMD, accounting for 32% of cases, has indicators pcews in one dataset and rews in another. Predicting the MPM from this scenario is 4.5 times more accurate on average than the worst-case scenario where all indicators come from separate datasets (p-c-r-e-w-s). The method closely predicts multidimensional poverty regardless of a country’s poverty level, not just on average.

Probability theory also allows us to calculate the absolute maximum and minimum possible overlaps between deprivations, given the fragmented data available. These bounds quantify uncertainty around any multidimensional metric predicted using the fusion method and can be used to correct for systematic bias. A simple correction using the median of these bounds substantially reduces prediction error.

Some final words

Together, these results demonstrate that fusing aggregated indicators can produce credible estimates of multidimensional poverty, particularly when datasets share overlapping indicators. This approach has the potential to expand global poverty monitoring — increasing coverage, incorporating dimensions such as health and security, and filling gaps in time series. We are working on applications to close these measurement gaps. This is not just a technical achievement: knowing not only how many people are poor, but how deprivations overlap and where, allows governments and development organizations to design more targeted and effective interventions.

READ MORE: https://blogs.worldbank.org/en/opendata/measuring-multidimensional-poverty-in-countries-with-incomplete-

You may also like