Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
Researchers have created a benchmark called TIDE to test how well large language models (LLMs) can understand documents that change over time. The benchmark includes 3,050 questions about customs instruments from the Government of Bangladesh between 1969 and 2025. Nine recent LLMs were tested on this dataset, but none performed exceptionally well, with a best accuracy of only 68.5%. The results suggest that LLMs are more likely to find correct versions than reject incorrect o
Researchers have created a benchmark called TIDE to test how well large language models (LLMs) can understand documents that change over time. The benchmark includes 3,050 questions about customs instruments from the Government of Bangladesh between 1969 and 2025. Nine recent LLMs were tested on this dataset, but none performed exceptionally well, with a best accuracy of only 68.5%. The results suggest that LLMs are more likely to find correct versions than reject incorrect ones, and tend to favor confident parametric answers over authoritative text.
---
Why it matters: This matters because it highlights the limitations of current large language models in handling temporal information, which is crucial for many real-world applications such as law, finance, and historical research. Improving LLMs' ability to understand evolving documents could have significant impacts on these fields.
Source: https://arxiv.org/abs/2608.08512
This article was originally published at: https://arxiv.org/abs/2608.08512