AI

Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal

Researchers have developed a method to improve the accuracy of quality appraisal in systematic reviews by analyzing disagreements between human experts and large language models. They used a checklist-based approach and found that certain items were more prone to disagreement, particularly those with ambiguous or conditional criteria. By revising these items, they improved agreement between humans and LLMs. The study suggests that analyzing human-LLM disagreement can help ide
Researchers have developed a method to improve the accuracy of quality appraisal in systematic reviews by analyzing disagreements between human experts and large language models. They used a checklist-based approach and found that certain items were more prone to disagreement, particularly those with ambiguous or conditional criteria. By revising these items, they improved agreement between humans and LLMs. The study suggests that analyzing human-LLM disagreement can help identify problematic checklist items and improve research synthesis workflows. --- Why it matters: This matters because it highlights the importance of checklist design in AI-assisted quality appraisal. Engineers and researchers working on natural language processing tasks will need to consider how their models interact with ambiguous or conditional criteria, and how this affects agreement with human judgments. Source: https://arxiv.org/abs/2608.20385

This article was originally published at: https://arxiv.org/abs/2608.20385