The Problem
Data Extraction Without Sovereignty
Language data gathered without consent, fair compensation, or provenance tracking. Communities have no visibility into how their data is used downstream.
Curation Without Community Voice
Annotation decisions made by external actors without consulting contributors. Once data enters a corpus, the trailback to the community disappears.
Model Training Concentrated in Few Hands
Compute infrastructure sits with well-resourced institutions. Low-resource languages are fine-tuned as afterthoughts, not designed in.
Deployment Skips Local Needs
Few AI use cases move from pilot to public service delivery at scale. AI deployments are frequently not solving user problems at the scale that matters — particularly when viewed through the lens of public interest AI, where success means citizens can access essential services, exercise rights, and participate in public life in their own language, not just that a product has been released. Access pricing often over-charges communities while delivering limited value.
No Value Return
No benefit-sharing mechanisms exist. Value concentrates at deployment in the hands of large actors. The loop never closes back to communities.
From the Ground: LLA Evidence
The Local Language Accelerator (LLA) is UNDP’s global programme supporting countries to develop inclusive language AI by strengthening the digitalization of low-resource languages. Working with governments, universities, technical partners, startups, and communities across 11 countries in Asia, Africa, Latin America, Europe, and the Arab region, the LLA supports the creation of language datasets, local AI ecosystems, and public-interest applications. Through two years of implementation, recurring bottlenecks have emerged across very different country contexts — in data collection and stewardship, model development, infrastructure, deployment, institutional uptake, financing, governance, and reuse. Here is what we are seeing in two active implementations.
Topics: ⚖️ Licensing & community rights; 🗣️ Dialect variation; ️ 🛡️Cross-border data protection; 🤝 Community trust
Topics: ⛔ Deployment wall; 🏛️ No government mandate; 📡 No value chain approach; 🌾 Agriculture advisory
Why This Matters for Development
The question of linguistic inclusion in AI is ultimately a question of who gets to participate in and benefit from the AI-driven transformation of economies and societies. AI has the potential to accelerate sustainable development, expand access to services and opportunities, and advance progress towards the Sustainable Development Goals. But its benefits will not be guaranteed. If the systems being built work better for some languages, communities and markets than others, AI risks reproducing - and potentially widening - existing inequalities.
This is particularly consequential for communities whose languages remain under-represented in digital systems. When AI does not work effectively in the languages people use to access healthcare, education, finance, public services and markets, linguistic exclusion can become a barrier to meaningful participation and impose a ceiling on the benefits gained from new innovations. Therefore, fostering linguistic diversity is a part of the broader challenge of ensuring that AI contributes to inclusive development and leaves no one behind.
Beyond Linguistic Diversity Toward Locally Sustainable AI
For UNDP, linguistic and cultural diversity in AI is not an end in and of itself. Rather, it is a condition for making AI work for development: for creating jobs, unlocking investment, building resilience, strengthening institutions, and enabling transparent, participative governance. AI is diffusing into economies and institutions, which signals that certain enabling conditions are already in place. Yet it is still uncertain whether diffusion is happening on terms that build agency: whether countries retain the capability to understand, negotiate, supervise and correct the systems entering their public and economic life, and whether public value can be built and retained locally rather than passed through or extracted out.
The choices that countries make now will shape who benefits from AI. They will influence jobs, investments, market resilience, public services and safeguards, as well as determine who creates value and who captures it. Both the Data-to-AI Value Chain and Every Language Matters are designed to equip government and communities with the analytical frameworks, policy tools, and diffusion pathways needed to shape, deploy, and benefit from AI on their own terms. This shifts the focus from whether a language can be represented in a model to whether the conditions exist for AI to become a meaningful and sustainable development tool for the people and communities it is intended to serve.