Index SLM Technical Report
Researchers at Bilibili have developed a series of open small language models called Index-1.9B. The four models in the series - Index-1.9B-Base, Index-1.9B-Pure, Index-1.9B-Chat, and Index-1.9B-Character - were pre-trained on 2.8 trillion tokens predominantly in Chinese and English. The team reports that the base model achieves competitive results with larger models on various benchmarks, including examination, reasoning, mathematics, and code tasks. The models are released
Researchers at Bilibili have developed a series of open small language models called Index-1.9B. The four models in the series - Index-1.9B-Base, Index-1.9B-Pure, Index-1.9B-Chat, and Index-1.9B-Character - were pre-trained on 2.8 trillion tokens predominantly in Chinese and English. The team reports that the base model achieves competitive results with larger models on various benchmarks, including examination, reasoning, mathematics, and code tasks. The models are released along with evaluation code on GitHub.
---
Why it matters: The development of Index-1.9B has implications for researchers working on small language models, as it provides a competitive alternative to larger models. Understanding the pre-training methods and architecture used in this project can inform the design of future small language models.
Source: https://arxiv.org/abs/2607.09885
This article was originally published at: https://arxiv.org/abs/2607.09885