AI

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

Researchers have created a new benchmark for evaluating the performance of retrieval-augmented generation (RAG) models on Canadian case law. The CanLegalRAGBench assesses how well these systems can retrieve relevant documents and generate accurate answers to legal questions, while also highlighting limitations such as hallucinations and overly detailed responses. The evaluation shows that open-source embedding models are competitive with closed-source models, but automatic ev
Researchers have created a new benchmark for evaluating the performance of retrieval-augmented generation (RAG) models on Canadian case law. The CanLegalRAGBench assesses how well these systems can retrieve relevant documents and generate accurate answers to legal questions, while also highlighting limitations such as hallucinations and overly detailed responses. The evaluation shows that open-source embedding models are competitive with closed-source models, but automatic evaluations may not accurately reflect a system's performance. --- Why it matters: This matters because it provides a more realistic and comprehensive way to evaluate the performance of RAG systems on Canadian case law, which is essential for ensuring justice and accuracy in legal decision-making. By identifying limitations and areas for improvement, this benchmark can help drive progress in developing more reliable and effective legal RAG systems. Source: https://arxiv.org/abs/2605.30497

This article was originally published at: https://arxiv.org/abs/2605.30497