AI’s Cultural Divide;

Are your Language and Culture Winning or Losing?

Introduction: Reframing the linguistic & cultural divide

The economic divide in AI access—rich versus poor, connected versus offline—is well documented. This paper identifies a second, deeper divide running underneath it: a linguistic and cultural gap that determines not whether populations can use AI, but how well it works for them, what it costs them in their own language, and whose worldview is reflected in the answers they receive.

Large language models are trained 80-90% on English data: GPT-3’s training was ~93% English; Llama 2 ~90%. Every other language was trained on whatever scraps remained. This has four cascading consequences: worse quality for non-English queries, higher costs due to tokenization design, exported Western cultural and values defaults, and the permanent absence of undigitized heritage from AI’s memory.

We have to remember that AI Models are a recent phenomenon that is still in its infancy. So flaws are expected and a normal part of developing them. The arguments made in this paper claim that the maybe AI models were not deliberately built bias and ignorant of other languages, cultures and values. AI corporates are working singularly and in cooperation with governments and NGOs and the United Nations to fix some of these issues.

Part I: The Mechanisms of the Divide

1. Training Data and the Quality Gap

English is spoken natively by roughly 5% of the world’s population yet constitutes close to half of all website content. G7 countries, ~8.5% of the global population, generate 70% of web content. Chinese, Arabic, and Hindi, spoken by nearly a quarter of the world’s population, generate less than 2% of the open web. To a model that trains on what is available on the web, the digital footprint is the only one that counts. So if you ask a frontier model a factual question in English, it performs near its best. Ask the identical question in an under-represented language in the web and accuracy drops. This is not a translation problem. The model genuinely knows less, because it read less of the language.

The Chinese exception: Even though its language makes up just 1.2% of the indexed web, yet over 1.1 billion internet users operate inside domestic platforms—WeChat, Baidu, Douyin—that sit in a version of the internet most AI models don’t see. The result is that Chinese models outperform Western ones on culturally specific benchmarks better suited for the Chinese culture. The same closed circuit that leaves China out of the open web allows a separate set of models to serve Chinese users unusually well.

2. The Token Tax: Cost as a Cultural Barrier

When you type a message into an AI model, the system chops text into tokens, and you are billed by the token (even if you’re using the free tiers). For English, this works cleanly: roughly four characters per token. However, when it’s asked in languages it was not designed for, it starts splitting words into far more pieces. In clear examples, Korean texts need 1.5 – 3×, Arabic texts need 3–4× more tokens than equivalent English texts. For other less spoken languages like Malayalam the cost may rise to 15.7 folds.

Three factors compound this:

  • Script shape: Non-Latin scripts (Arabic, Ethiopic, Burmese) take more digital storage per character, even before tokenization starts.
  • Linguistic structure: Some languages pack grammar into the word itself where English adds separate words—penalizing even some Latin-script languages.
  • Neglect: Tokenizers learn shortcuts from whatever text they were trained on. Therefore they have efficient shortcuts for English and almost none for anything else.

The languages that end up at the very worst end of this tax, are the ones unlucky enough to get hit by all three at once.

The implication is that a $2 English query can become a $9, $15, or even $20 query in other languages—for the exact same words. This is not a market failure; it is a structural feature of how the technology was built. The closer your language is to Latin, the cheaper AI is for all your users: students, researchers, and investors.

3. Cultural Defaults: Whose Values Ride Along

Everything so far has been about which languages made it into the data. There’s another divide related to which humans made it into the room where decisions got made.

We found that the US and China produce 96% of the top 50 frontier AI models (48 models out of 50). Inside those labs: only 12–18% of AI researchers are women; just 14% of research papers have a woman as first author. At major US tech firms, white and Asian employees hold 85–87% of technical roles; Black employees make up roughly 3% of Silicon Valley’s technical workforce.

This is a product defect. A 2026 study of an AI tutoring tool built for Indian classrooms asked for examples of vitamins in a typical meal. It suggested Caesar salad, because the model was trained on Western data by people who had never thought about what “a typical meal” looks like outside their own kitchen.

It isn’t just about vocabulary — a model can speak fluent Arabic, Hindi, or Russian and still reason with an American’s assumptions about family, gender, what counts as a reasonable work-life balance, what polite sounds like, even what a healthy meal is. And once you notice that pattern, the outputs themselves start making more sense.

4. AI Gender, Race, Language, and Religious Biases:

AI models have also performed other biases, some of which tremendous effort has been spent worldwide and over decades to eradicate. Unfortunately some of them still found their way into the latest 21st century technology, threatening to bring back painful chapters of the world’s joint past. These are some of the most outstanding biases we found.

Gender: A 2025 Nature study found a consistent pattern where women get portrayed as younger than men in the same roles. The gap is widest in high-status jobs specifically. When researchers had ChatGPT generate résumés, it made the women 1.6 years younger and less experienced on average, and rated the older male applicants as more qualified. Ask an image model for a “manager” and an “assistant” side by side, and it reliably hands masculine traits to the manager and feminine traits to the assistant, unprompted.

Geography and Race work in a similar way. Another study (proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency) found that asking for a poor person for example, overwhelmingly generates black faces (has been changed since to non-white faces). Another study (Naik,R., & Nushi, B. 2023) measured how far AI-generated images of everyday life deviate from the model’s unprompted default, country by country: the USA, Australia, and Germany sat closest to default. Nigeria, Ethiopia, and Papua New Guinea, sat furthest away. Those countries get treated as the exception needing special handling, while the West is treated as simply normal.

Religion follows a similar pattern. Researchers testing GPT-3 back in 2021 found 66% of completions involving Muslims spontaneously mentioned violence or terrorism, versus 5-15% for every other major religion tested. Later research (Hemmatian,B., Baltaji, R., &Varshney, L.R.2024) found that the gap persisted even after developers specifically tried to fix it, citing that Open AI had made efforts to reduce the original bias. They also found that the association with violence persisted-just in subtler, less obvious forms.

Even happy family and a good life aren’t neutral concepts to a model. Stanford researchers tested the literal phrase -a happy family- and got a small, two-parent household by default, not the multi-generational or extended-family households that are the actual lived norm across much of Asia, Africa, and the Middle East.

The academic term for this underlying pattern is WEIRD: Western, Educated, Industrialized, Rich, and Democratic. It describes cultures that tend to moralize around individual rights and personal choice, versus the majority of the world’s cultures, which moralize more around communal duty and obligation. It’s not that one is right and one is wrong. It’s that a model trained mostly on WEIRD text picks one silently and hands it back as if it were the only option.

As pointed out above, some of this is already shifting. It has been noted that Open AI and Microsoft’s Bing now quietly inject demographic diversity into prompts that don’t specify one. However, Google had to publicly roll back a 2024 version of this same approach after it overcorrected, generating historically inaccurate images. Furthermore, a 2025 study found a genuine trade-off between diversity and factual accuracy that researchers are still working through. The underlying training –data bias these patches sit on top hasn’t gone away; it’s being covered till they are actually solved

5- The vanishing Heritage problem

Another important gap that may not fully close is related to whole heritages being forgotten. AI models learn from what is on the internet. The internet which is barely three decades old represents a thin sliver of human history. Centuries of literature, record-keeping, and oral tradition predate it. For most cultures, that material sits in physical archives or in memory. Nobody with funding or institutional mandate got around to scanning, transcribing, and uploading it, and so AI models simply don’t see them. For AI they don’t exist.

Again China is an exception. It has realized this gap, and started work on fixing it, through a joint project between the private sector and the public sector: ByteDance — the company behind TikTok — teamed up with Peking University to build a platform where AI does the first pass on scanned ancient texts, then a volunteer community, checks the machine’s work by hand. In 2025 alone, those volunteers proofread 1.5 billion characters across roughly 20,000 ancient books. They all decided, deliberately, that centuries-old heritage was worth the money and the manpower to save.

In other cultures entire local histories are kept only in fragile paper. Very few got the Chinese treatment. Today, it’s not just that some languages are underrepresented online today. It’s that some cultures had the resources to move their past onto the internet before AI came along to learn from it, and most didn’t — which means the users of AI systems being built right now are witnessing which parts of human history get to exist in the age of artificial intelligence and which parts simply don’t. What is more unsettling is that most of what is disappearing left no digital trace.

6- Winners, Losers, and the K-Shape

Economists use a K-shape to describe post-pandemic recovery: one arm goes up (winners), the other down (losers). AI’s K is murkier.

On the top arm: English speakers and well-resourced languages, who get the best model quality, pay the lowest token cost, and see their own cultural assumptions reflected back as the default. Add China’s own users, who benefit from their closed domestic ecosystem.

On the bottom arm: speakers of the world’s under-resourced languages, who get worse answers, pay several times more in tokens for the privilege, and receive a worldview from a culture that isn’t theirs.

The Case for A Murky K

The K is not so clear cut. There’s more to the story.

The Invisible Group: beneath even the bottom K arm sits an invisible group with no arm on the K at all — the cultures whose history and oral tradition never made it online in the first place, and so don’t get an outcome from AI.

India’s complicated situation: India is the world’s largest source of AI training labour. It has more than 500,000 workers in this field. However, it doesn’t own a single frontier model. And even though part of its labourers work is directed towards Hindi, Tamil, and other Indian languages, a large share is still general-purpose English-content work. This makes India in a category of its own: not a clean winner (no model ownership, no proportional pay for the work: the average pay is $400), but not a clean loser either (steady, growing income, real technical skill-building)

So the K isn’t just about who has better model access or better model quality. It’s about whose history, heritage, language, and values make it into the future, and those who got forgotten and left behind. It’s also about the workforce supplying the raw material of intelligence to both the American and Chinese poles, and getting paid a rounding error of what that intelligence is worth once it’s built.

Part II- Solutions and implications

Many have noticed these divides (including the corporates responsible for frontier AI models) and are doing something about it. Most are in their early stages, and some are actually succeeding. This has important implications for all stakeholders, and may provide profit incentives for investors to back these efforts.

1- Who’s Already Trying to fix the divide:

Sovereign AI:

Many governments have already realized these facts and are working on finding the solutions that best fit their countries through what’s called Sovereign AI. The term itself refers to a nation’s strategic effort to develop and control its own artificial intelligence infrastructure; including data, models, and computing power, rather than relying entirely on foreign companies or governments.

For them, it’s about digital self-determination. Some of the most outstanding examples are the Mistral project in France, and to a lesser extent: the Jais and Falcon projects in the UAE, and the Indian Bhashini project.

Countries are racing to build sovereign AI for 3 main reasons:

a. Data Sovereignty, Security and keeping data local: Countries worry that foreign AI systems could expose sensitive government, military, or industrial secrets to rivals, and they would like to keep data in accordance with their local regulations and laws.

b. Economic & Geopolitical Independence. Most countries wish to avoid ‘AI colonialism’ through avoiding paying perpetual licensing fees to foreign tech giants for critical infrastructure. Also they seek avoiding being geopolitically vulnerable.

c. Cultural & Linguistic Preservation and Preventing cultural erosion and Language survival. Sovereign AI ensures their native language, idioms, and traditions are properly represented.

This is domestically. On the international level there is a growing and significant movement toward internationally inclusive AI data models. This movement consists of numerous collective efforts from governments, NGOs, private companies, and the corporates owning AI models.

Here are some key examples of this work in action:

i-The UNESCO published earlier this year the Global Roadmap on Multilingualism in the Digital Era that pushes for something subtler than just ‘make more dataset’”. It argues that the communities that speak different languages should actually own and govern the data being collected about them, rather than having it collected, used, and monetized by outside institutions with no obligation to give anything back. This is a major policy step.

iiMozilla’s Data Collective: A community-owned data market to ‘build an equitable AI ecosystem rooted in community, choice, and autonomy… and where Communities can sell their datasets, define the terms of use, and have a direct input on which initiatives aligned with their values should benefit from their data’. This is what their site says.

iiiMicrosoft’s AI for Good Lab: whose work includes ‘preserving cultural heritage with AI by digitizing historic sites, safeguarding endangered languages and expanding access to cultural knowledge’. It is also working on the LINGUA Africa project.

ivMasakhane and LINGUA Africa: A Pan-African initiative that started with researchers simply refusing to accept that African languages would stay an afterthought. It’s grown enough now that Microsoft’s AI for Good Lab, Google.org, and the Gates Foundation have all joined in through a joint funding effort called LINGUA Africa.

v- Global RETFound Consortium: A medical effort involving over 100 research groups from 65 countries to build a global medical AI. Trained on over 100 million eye images, it aims to create an AI that is geographically and ethnically diverse.

vi- The United Nations Development Programme (UNDP) language initiatives, launched in April 2026, and ‘focus on closing the artificial intelligence (AI) language gap for underrepresented and low-resource languages through digital tools, local accelerators, and inclusive policies

2-Policy & Business Implications

We have to remember that AI Models are a recent phenomenon that is still in its infancy. So flaws are expected and a normal part of developing them. The arguments made in this paper claim that the maybe AI models were not deliberately built bias and ignorant of other languages, cultures and values. AI corporates are working singularly and in cooperation with governments and NGOs and the United Nations to fix some of these issues.

However, there is so much work left to be done in this respect; whether it be in raising awareness, or actually fixing them. Much of it can be both profitable and beneficial nationally, regionally, and internationally. Let’s take a look at the implications for each main group involved:

a- For governments, the lesson is that waiting for a foreign lab to prioritize your language is a losing strategy. The countries making real progress didn’t wait for permission; they funded their own datasets, their own tokenizers, and their own small teams, often with small budgets. That’s a policy choice available to almost any government. The China case adds a sharper lesson still: heritage digitization is a strategic decision about whether your civilization exists inside the tools shaping the next century of information, or not.

bFor Corporate businesses and Investors: The token tax and the quality gap are current, exploitable pricing failures. “Is our AI actually good in this market?” is now a legitimate line item in market entry planning, and public policy alike, not an afterthought to raise once the English version already works. Every business serving Arabic, Hindi, or Malayalam speaking markets today is quietly eating a cost and quality penalty that most of their competitors serving English-speaking markets never notice. Here is where an actual opportunity sits: not in building another general-purpose English model, but in building for the underserved: better tokenizers, better regional datasets, and better sovereignty-flavoured models for markets nobody else is fixing. The market exists for this.

cFor NGOs and grassroots: Previous initiatives like the pan-African initiative Masakhane, prove that will and capability are pre-requisite to funding, not the other way around. The initiative worked several years successfully before corporate money came in. As for grassroots, they have a role in raising awareness about these issues from the ground, from the communities that are facing them. Otherwise their communities’ heritage may become forgotten too.

Conclusion: Is it too late to dream about a World Wide AI?

The internet was supposed to be the World Wide Web — one shared space, open to anyone with a connection. It never quite delivered on that promise. AI showed up promising something bigger still: a shared intelligence, fluent in anyone’s question regardless of who they are or where they’re asking from. So, can we still get to dream of a World Wide AI, or is it already too late for that?

Left alone, the market builds for the biggest, richest, best-documented audience first, which meant English, then Chinese, then whoever was profitable enough to bother next. Nobody sat down to erase the Shan language or leave West African traditions outside the training data. It just wasn’t anyone’s job to stop it. No one set out to draw a non-white as a poor man, or a woman less qualified than a man. It’s what the ‘web-crawlers’ found.

However, this doesn’t have to be the case. We have learnt our lesson of where this ignorance leads. So many countries, international organizations and grassroots NGOs are taking the initiative to correct this. A pan-african initiative building 400 models with almost no budget, and a government, academia, and tech-company partnership in China proves that centuries of heritage can be saved if someone simply decides it’s worth paying for. None of that happened by default. Every single fix in this paper exists because a specific group of people raised the issue and refused to wait for someone else to prioritize them.

A World Wide AI isn’t coming automatically, and it isn’t gone forever yet. It’s a choice, to be made over and over, by whoever funds the unglamorous work of including one more language, one more archive, and one more community that the default settings were never built for. The dream for a worldwide AI is just waiting on more people to decide it’s worth the trouble. And for investors and businesses to realize that this can also be profitable.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *