AI

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

Researchers have created a new benchmark called MGAL to evaluate the performance of large language models across different languages and levels of granularity. The benchmark consists of United Nations reports in six official languages, with varying lengths from 8K to 128K tokens. It assesses models' ability to comprehend text at word, sentence, paragraph, and document levels, as well as their position within the document. Experiments showed that large language models perform
Researchers have created a new benchmark called MGAL to evaluate the performance of large language models across different languages and levels of granularity. The benchmark consists of United Nations reports in six official languages, with varying lengths from 8K to 128K tokens. It assesses models' ability to comprehend text at word, sentence, paragraph, and document levels, as well as their position within the document. Experiments showed that large language models perform well on fine-grained tasks but struggle with coarser ones, and closed-source models have an advantage in lower-resource languages. --- Why it matters: This matters because it provides a more comprehensive evaluation of long-context comprehension across different languages and granularities, which is essential for developing more robust and accurate language models. It also highlights the challenges faced by current models in handling coarser-grained tasks and local semantic crowding. Source: https://arxiv.org/abs/2608.20853

This article was originally published at: https://arxiv.org/abs/2608.20853