Artificial intelligence (AI) is often described as a general-purpose technology (GPT) — an innovation so profound and universal that it fundamentally alters entire economies and societies. As a universal tool, AI has the potential to bridge gaps in human understanding and democratize access to knowledge.
Much of AI’s utility stems from its ability to predict and generate human-like text by analyzing vast datasets gathered from across the internet. Its capabilities are largely dependent on this foundation. What happens when that data is uneven? What if certain voices and information are far harder to come by than others?
English alone accounts for nearly half of global URLs, and languages spoken in wealthy countries are represented extensively across the web. Languages spoken primarily in low and middle-income countries, however, are underrepresented–even when they have hundreds of millions of speakers. If the data aren’t there for the AI models to train on, they will be less powerful, less useful, and AI will be unable to become a truly universal GPT.
To get a sense of this unevenness of representation, we can visualize the relative prevalence of different languages by plotting the number of speakers against the number of URLs in that language.
As a result, AI models often lack the accuracy and nuance to be useful where they could matter most, in places where knowledge and information is hard to access. Closing this gap will require deliberate investment, including expanding text data in underrepresented languages and supporting local AI development capacity.
AI can become a truly universal technology, but it requires not leaving the languages often spoken in poorer regions of the world behind.
For more information, check out the Atlas of Global Development, which dives deeper into global inequities in the use of AI.
read more: https://blogs.worldbank.org/en/opendata/how-well-does-ai-speak-your-language-