Can AI Decide Which Languages Matter?
- Giulia Tricamo
- May 14
- 8 min read
Updated: 11 hours ago
I don’t really think about what language AI speaks.
I just type.
Sometimes I write in English, sometimes in Italian, and sometimes I start a sentence in one language and finish it in another without even noticing. If I’m trying to explain something complicated, I might use English because the words come more easily. If I’m looking for a particular expression or trying to describe something that exists in my head in Italian, I’ll switch back without thinking twice.
I’ve always assumed that if AI can understand both languages, then, basically, the languages are equal.
But is that actually true? Does AI understand every language equally well?
Apparently, not quite.
I actually asked ChatGPT this question because I was curious. I wasn’t expecting it to turn into a whole topic, but the answer mentioned the difference between languages with lots of digital resources and those with much less representation online. That led me to an article in Frontiers in Psychology about what the authors call “algorithmic dominance,” and suddenly I was thinking about language in a completely different way.
Some languages are much better represented in the data AI learns from than others, and that difference can affect how well these systems work.
That sounds like a technical problem, but the more I thought about it, the less technical it seemed.
AI doesn’t understand every language in the same way
One of the strange things about AI is how neutral it feels when you use it. You type something, it responds, and there is nothing obvious in between to remind you that the system might be working very differently depending on what language you are using.
But AI systems don’t learn languages in the same way that people do. Large language models are trained on enormous amounts of data, and the amount of digital material available in each language varies enormously. Some languages have huge numbers of books, websites, articles and other online resources, while others have a much smaller digital presence.
That difference matters.
A 2026 study presented at the Association for Computational Linguistics conference tested multilingual large language models across 61 languages using 3.9 million samples. The researchers found a persistent performance gap between English and lower-resource languages, even among models designed to work across many languages.
The basic idea is actually quite easy to imagine. If two people ask an AI exactly the same question in two different languages, one might receive a detailed, natural answer while the other gets something awkward, incomplete or slightly inaccurate. Neither person has done anything differently. Their languages simply don’t have the same level of support behind them.
And suddenly, the technology doesn’t feel quite so neutral.
But why should that matter?
At first, I thought this was mainly a technology problem. If AI is worse at a particular language, then developers should improve it. Problem solved.
But the more I thought about it, the more I realized that language is about much more than whether a computer can produce a grammatically correct sentence.
Language carries jokes that don’t always work in translation, expressions that don’t have an exact equivalent, family habits, memories and references that might make sense immediately to one group of people and not at all to another. Even the way we tell a story can change depending on the language we are using.
So when a language is poorly represented online, it isn’t simply that a computer has fewer words to work with. There may also be less digital representation of a particular way of expressing ideas, humor and experiences.
That made me wonder whether being represented digitally could eventually become part of what it means for a language to have a place in the modern world.
We already spend so much of our lives online. We communicate, study, search for information, watch videos, write, create and increasingly interact with technology through digital platforms. If a language isn’t well supported in that environment, does that eventually affect the people who speak it?
I don’t know, but I think it’s a question worth asking.
English already has an enormous advantage
This isn’t really a problem that AI created.
English has had an enormous amount of global power for a long time. It is dominant in many areas of science, business, international education, technology and popular culture, so it isn’t surprising that the internet contains an enormous amount of information in English.
AI is being built on top of that existing world.
If most of the information available to a model is in English, then English becomes easier for the system to process. If the technology works particularly well in English, people have even more reason to use it. And if people use it more, even more information gets produced in English.
That can create a cycle: more data leads to better technology, better technology encourages more use, and more use creates even more data.
Meanwhile, languages with less digital representation may have a harder time catching up.
The research I found doesn’t mean that this cycle is inevitable, but it made me realize that technological inequality doesn’t necessarily begin with a deliberate decision to favor one language. Sometimes it can simply grow out of the world that already exists.
Then there is the psychological side
This is where the question became more interesting to me because it connects technology with something much more personal.
What happens when you are constantly reminded that one of your languages works better than another?
A 2026 article in Frontiers in Psychology explores the idea of “algorithmic dominance” and argues that unequal digital support for different languages could potentially affect how people perceive their own languages and their ability to use them in digital spaces. The authors discuss possible connections with things like language anxiety, confidence and willingness to communicate, while making it clear that these are areas for future research rather than established causal effects.
I think that distinction is important, because it would be easy to take the idea too far.
But even as a possibility, it made me think.
Imagine being a student who is much more comfortable speaking a language that isn’t the dominant language online. You ask an AI to help you write something, and the answer is simply better when you use English. The first time, you might not think much about it. The second time, you might do the same thing because it’s easier.
Eventually, you might stop asking whether your first language would have worked just as well.
No one has told you that English is better. The technology has simply made one choice easier than the other.
That kind of influence is much harder to notice because it doesn’t feel like pressure.
It just feels convenient.
But AI could also help
This is where I don’t think the story should become simply “AI is bad for languages.”
In fact, AI could become an incredible tool for languages that have historically received less technological support. It could help document languages, make translation more accessible, help people learn languages that aren’t widely taught, and make old texts and other linguistic material easier to access and preserve.
Researchers are already working on ways to evaluate and improve multilingual AI systems rather than simply accepting English as the standard. So AI could reinforce existing language inequalities, but it could also help reduce them.
That’s what makes the question more interesting to me.
The problem isn’t necessarily that AI speaks English. The problem is what happens if we build a digital world where some languages are consistently easier for technology to understand, search, translate and reproduce than others.
If AI becomes one of the main ways we access information and communicate with each other, the languages it understands best could end up having an advantage that goes beyond technology.
Who decides what makes a language important?
This is probably the question I keep coming back to.
We often talk about languages as if some are naturally more important than others. English is a global language, some languages are considered useful for business or international communication, while others are described as rare or endangered.
But useful to whom?
And important according to what?
There are thousands of languages spoken around the world, many of them by relatively small communities. The number of people who speak a language doesn’t necessarily tell us how important that language is to the people who use it, or how much history, culture and identity it carries.
A language can contain stories, traditions, relationships and ways of understanding the world that can’t really be measured by the number of websites or documents available in it.
Digital systems, however, need data, and data is much easier to measure. We can count how many websites exist in a language, how many people use it online, how many books and documents are available, or how accurately an AI model can process it.
Those numbers can start influencing which languages receive more attention and technological investment, even though they tell us very little about their cultural value.
This is where technology starts becoming cultural. Not because an AI has consciously decided that one language or culture matters more than another, but because the infrastructure behind it can reflect the inequalities that already exist in the world.
And human societies are not neutral.
They have histories, inequalities, dominant languages and minority languages. AI doesn’t start from zero; it starts with everything we have already created.
I don’t think the answer is to reject AI
I use AI, and I find it incredibly useful. It has made communication across languages easier and has the potential to do something really positive for languages that have historically received less technological support.
That’s actually part of why I find this issue so interesting. I don’t think technology has to be deliberately biased for its effects to be unequal. Sometimes it is enough for a technology to work better for some people than others.
We tend to think of algorithms as objective because they are machines, but machines are trained on information created by human societies. They inherit some of the patterns, assumptions and inequalities that already exist.
That doesn’t mean AI is destined to reproduce them forever. It means that we have to notice them if we want to change them.
So, can AI decide which languages matter?
Not by itself. At least, I don’t think it can.
But it can influence which languages are easier to use, easier to search, easier to translate and easier to participate in online. Once something becomes easier to use, people naturally tend to use it more, and that can create its own kind of power.
Maybe the biggest question isn’t whether AI will make us stop speaking certain languages. Maybe it is whether we will notice the hierarchy being built around those languages before we start treating it as normal.
If technology becomes one of the main places where we communicate, learn and create culture, then being represented there matters. A language shouldn’t have to become economically powerful before it deserves to be technologically visible.
And maybe the future of multilingualism isn’t simply about keeping languages alive in the physical world. It might also mean making sure that people can bring those languages with them into the digital one.
I started thinking about this because I wondered whether AI understood English and Italian equally well.
I ended up wondering something much bigger:
If technology becomes part of the way cultures survive, who gets to decide which cultures are heard?
I don’t have a complete answer yet. But I think that’s probably a better place to start than assuming technology is neutral simply because it is a machine.
Sources:
Jianbu Yang, Jie Zeng and Xingpeng Zheng, From linguistic imperialism to algorithmic dominance: psychological implications of language hegemony in the age of digital intelligence, Frontiers in Psychology, 2026.
Wenhan Han et al., MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages, Findings of ACL 2026.



Comments