NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
Researchers have developed a multilingual language model called NE-BERT that can process nine Northeast Indian languages, including Hindi and English as anchor languages. The model was trained on approximately 8.3 million sentences and outperforms existing models in these languages. It also addresses vocabulary fragmentation issues in extremely low-resource languages like Pnar and Kokborok through aggressive upsampling strategies.
Researchers have developed a multilingual language model called NE-BERT that can process nine Northeast Indian languages, including Hindi and English as anchor languages. The model was trained on approximately 8.3 million sentences and outperforms existing models in these languages. It also addresses vocabulary fragmentation issues in extremely low-resource languages like Pnar and Kokborok through aggressive upsampling strategies.
---
Why it matters: This matters to AI researchers because it provides a more inclusive language model that can handle diverse languages, which is crucial for natural language processing tasks such as part-of-speech tagging. This could also support digital inclusion efforts in Northeast Indian communities by providing them with more accessible and accurate language tools.
Source: https://arxiv.org/abs/2608.18094
This article was originally published at: https://arxiv.org/abs/2608.18094