Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models
A new benchmark called Know2Guess has been developed to evaluate the reliability of large language models. The benchmark separates supported answering from unsupported guessing and data contamination, which can be a problem in current evaluation methods. It contains 1,200 items across five domains and provides metadata about potential contamination risks. Researchers have tested several popular models using this benchmark and found that while some models perform better than o
A new benchmark called Know2Guess has been developed to evaluate the reliability of large language models. The benchmark separates supported answering from unsupported guessing and data contamination, which can be a problem in current evaluation methods. It contains 1,200 items across five domains and provides metadata about potential contamination risks. Researchers have tested several popular models using this benchmark and found that while some models perform better than others, none of them are perfect. The dataset is publicly available for further testing.
---
Why it matters: This matters to researchers in AI because it provides a more accurate way to evaluate the reliability of large language models, which can be prone to contamination risks. By separating supported answering from unsupported guessing and data contamination, this benchmark helps identify areas where models are weak or biased.
Source: https://arxiv.org/abs/2606.26101
This article was originally published at: https://arxiv.org/abs/2606.26101