Poverty is multidimensional, but measuring it that way is not easy. Data on some dimensions may not be collected in the first place due to budget constraints, conflict, or other barriers with no immediate solution. Or the data may exist but are fragmented across several datasets that are difficult to combine. For instance, we can estimate the share of low-income households from one dataset, those lacking basic infrastructure from another, and those with low education from a third – yet no single source captures all three dimensions together, making it difficult to know how deprivations overlap. This fragmentation is not just a technical inconvenience, it shapes what governments see, and therefore what they act on. Fragmented data leads to fragmented policies.
In what follows, we describe the fragmented state of poverty data and the cost of leaving those gaps unaddressed. Our companion blog outlines a method for combining datasets to measure multidimensional poverty in these scenarios.
The fragmented landscape of poverty data
While a wealth of household survey data exists for many countries, these datasets typically live in parallel, non-overlapping universes. For example, the World Bank’s Global Monitoring Database (GMD) captures income/consumption and basic demographics across more than 160 economies, while Demographic and Health Surveys (DHS) capture demographic and health information of women and men of reproductive age across more than 90 economies. Accessing different data sources is costly, and even with access, combining them is not straightforward. They almost never sample the same households, making it impossible to link individuals across surveys.
A recent paper documents the gap in multidimensional poverty data, finding that only two indicators – educational attainment and phone ownership – are collected across four major cross-national survey databases: EU statistics on income and living conditions (EU-SILC), Gallup World Poll, Demographic and Health Surveys/Multiple Indicator Cluster Surveys (DHS/MICS), and Socio-Economic Database for Latin America and the Caribbean (SEDLAC). Figure 1, adapted from this paper and extended to include the World Bank’s GMD, illustrates how the global data landscape operates in thematic silos. To take one example, SDG 7 calls for universal access to electricity. However, the data on electricity access is available in four of the five databases. And where it is available, the overlap between lack of electricity access and any other deprivation can be observed for only 61 countries in Latin America and covered by EU-SILC, representing less than 8 percent of the global population.
read more: https://blogs.worldbank.org/en/opendata/beyond-the-dashboard–the-high-cost-of-fragmented-poverty-data