Lexicography is entering a breakthrough era. In March 2026, Google DeepMind unveiled 'LexiGen,' a specialized language model capable of generating complete dictionary entries—definitions, etymologies, usage examples, registers—for any language from a corpus of just 10,000 sentences. The technical feat relies on multilingual transfer learning trained on 400 languages and refined by native linguists. The project, dubbed the 'Universal Dictionary Initiative,' aims to document the 7,168 living languages listed by Ethnologue by 2030.
The urgency is real. According to UNESCO, a language disappears every two weeks. Of the 7,168 languages spoken worldwide, 3,045 are considered endangered, and 573 are 'critically endangered'—meaning only a few elderly speakers still master them. Until now, documenting a language required years of fieldwork by specialized linguists. With LexiGen, the bootstrapping phase—creating a basic vocabulary of 15,000 entries—can be completed in six weeks, provided minimal audio recordings and transcriptions are available.
The Académie française, an institution traditionally wary of technology, has surprised by announcing a partnership with the MIT Media Lab to integrate AI into the ninth edition of its Dictionary. The project does not aim to replace academicians but to equip them: AI analyzes 28 billion words from contemporary Francophone corpora—social networks, podcasts, scientific publications—to identify neologisms, semantic evolutions, and loanwords. In 2025, the system detected 4,200 new usages, 340 of which were retained by the Dictionary Commission.
Commercial applications are multiplying. Berlin-based startup Langify, which raised 85 million euros in funding in January 2026, offers companies real-time multilingual terminology services: a technical term created in English is instantly translated, defined, and contextualized in 120 languages with a claimed accuracy of over 97%. Airbus, Siemens, and the WHO are among its first clients. The model directly threatens traditional bilingual dictionary publishers, whose revenue has fallen by 45% since 2020.
The ethical question remains: who defines a language? Dictionaries have always been instruments of power, codifying certain usages and excluding others. Entrusting this codification to algorithms trained on massive corpora risks reinforcing dominant biases—English still represents 56% of global textual data. Linguists from the Endangered Languages Project emphasize the need for community governance: it is the native speakers, not Silicon Valley engineers, who should validate dictionary entries. AI is a tool, not an oracle—and linguistic diversity is too precious a heritage to be entrusted solely to algorithms.



