Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
Researchers have developed a new language model called VEXMLM to improve the performance of multilingual AI systems on low-resource Ge'ez-script languages such as Amharic and Tigrinya. The model expands the vocabulary of existing XLM-R models by adding 30,000 Ge'ez-script subwords and uses custom tokenizers trained on curated monolingual corpora. In experiments, VEXMLM outperformed other models on tasks like question answering and named entity recognition for Amharic and Tigr
Researchers have developed a new language model called VEXMLM to improve the performance of multilingual AI systems on low-resource Ge'ez-script languages such as Amharic and Tigrinya. The model expands the vocabulary of existing XLM-R models by adding 30,000 Ge'ez-script subwords and uses custom tokenizers trained on curated monolingual corpora. In experiments, VEXMLM outperformed other models on tasks like question answering and named entity recognition for Amharic and Tigrinya, with improvements also seen in 17 other African languages.
---
Why it matters: This matters to researchers working on AI for low-resource languages because it shows that vocabulary expansion and tokenizer adaptation can be a more efficient way to improve model performance than retraining from scratch. The results have implications for the development of multilingual AI systems that can handle underrepresented languages.
Source: https://arxiv.org/abs/2607.15209
This article was originally published at: https://arxiv.org/abs/2607.15209