GDF.Life

Languages without digital borders: how AI helps preserve linguistic diversity

Practical Workshop. Language, Culture, Data: AI for Low-Resource Languages. © RIA Novosti / Grigory Sysoev
Practical Workshop. Language, Culture, Data: AI for Low-Resource Languages. © RIA Novosti / Grigory Sysoev

Participants in the session, Language, Culture, Data: AI for Low-Resource Languages, discussed how artificial intelligence can become a powerful tool for preserving and developing the linguistic and cultural heritage of different peoples.

“UNESCO, as a leading UN institution, has identified language digitalization as a key area of work. Global experience shows that artificial intelligence is already becoming a powerful tool for supporting languages,” said Feride Aroniya, moderator of the session and Head of the Language Projects Support and Implementation Department at the House of the Peoples of Russia, setting the tone for the discussion.

President of the Gambella University Diriba Eticha Tujuba described Ethiopia’s experience in preserving low-resource languages through digital technologies and artificial intelligence.

He said that, rather than developing new AI models from scratch, it was important to draw on existing research, archives, and the knowledge of native speakers. He also emphasized the important role universities play in language preservation.

“Universities can process, validate, and document the resources needed to develop AI and work with these languages,” he said.

For example, Gambella University staff work with 84 languages, more than 25 of which have already been digitized and documented. These resources are available on a national platform and can be used to develop language technologies, educational programs, and mobile applications.

Director General of the Office of Public Radio and Television of Madagascar Festin Elisee Lemana noted that low-resource languages are not necessarily endangered or spoken by small populations. He cited Malagasy, the language of Madagascar, as an example.

“We are not trying to save endangered languages. We are trying, at the very least, to prevent a living, widely spoken language from being poorly represented in the new digital world,” he said.

One way to address this problem is to digitize audio archives and printed texts and bring them together in a unified database. Lemana believes that this approach will help develop AI tools for languages, automate the transcription of recordings, and preserve linguistic features.

Technical Director of the Languages of the Peoples of Russia Project at Yandex Andrei Mikheyev discussed the project’s efforts to expand the presence of national languages in the digital environment.

He emphasized that even widely used languages can remain low-resource for AI because of a lack of data in machine-readable formats.

Mikheyev identified the illusion of digitization as one of the key problems. This occurs when scanned books and archives are accessible to people but cannot be used to train AI models. Other obstacles include copyright restrictions, a lack of digital skills among native speakers, and insufficient funding for language-preservation projects.

Head of the Digital Ecosystem Agency at the Indonesian Chamber of Commerce and Industry Firlie Hanggodo Ganinduto discussed how an AI system’s ability to communicate in a local language does not guarantee that its responses will be safe or reliable.

In his view, users of low-resource languages may be particularly vulnerable because fraud detection and content moderation systems tend to work less effectively in these languages.

Devi Bahara Rachmawati, a lecturer at Universitas Indonesia, stressed that AI development depends not only on computing power and large-scale models, but also on an understanding of local languages and cultural contexts.

Mikhail Anisimov, Adviser to the Director for External Communications at the .RU/.RF Domain Coordination Center, discussed the development of technical support for national languages on the internet.

“Languages appear in many different forms on the internet and in the digital environment. Previous speakers mainly discussed language as content – as the substantive part of websites, resources, and the sources we work with. But language also features in various technical elements, such as domain names. Since 2010, it has been possible to create internet addresses using non-Latin scripts,” he said, explaining that this had expanded the possibilities for using national languages in domain names and email addresses.

Anisimov also highlighted international cooperation with UNESCO and other organizations. He noted that preserving languages in the digital environment also requires technical solutions that enable different languages to be used online for communication, service delivery, and the development of new technologies.

Other news
One of the main challenges for the media is to reconcile technological development with audience trust
One of the main challenges for the media is to reconcile technological development with audience trust
GDF Participants Discuss the Role of AI in City Management
GDF Participants Discuss the Role of AI in City Management
Boris Vasilyev: ITU Should Become the Main Forum for Developing Measures to Counter Starlink
Boris Vasilyev: ITU Should Become the Main Forum for Developing Measures to Counter Starlink
Cosmonaut Sergei Prokopyev: There are work chats on the International Space Station
Cosmonaut Sergei Prokopyev: There are work chats on the International Space Station
All news