AI

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

Researchers conducted a multi-site study to evaluate the performance of large language models (LLMs) in abstracting data from electronic medical records for clinical registry abstraction. They found that LLMs were less accurate than human abstractors, with accuracy declining as the ambiguity and required clinical reasoning increased. The study used a taxonomy of six categories to classify questions by their level of ambiguity and clinical reasoning.
Researchers conducted a multi-site study to evaluate the performance of large language models (LLMs) in abstracting data from electronic medical records for clinical registry abstraction. They found that LLMs were less accurate than human abstractors, with accuracy declining as the ambiguity and required clinical reasoning increased. The study used a taxonomy of six categories to classify questions by their level of ambiguity and clinical reasoning. --- Why it matters: This matters because it highlights the limitations of LLMs in complex tasks like clinical registry abstraction, which requires nuanced understanding of medical concepts and context. This has implications for researchers developing and evaluating LLMs for healthcare applications. Source: https://arxiv.org/abs/2608.20373

This article was originally published at: https://arxiv.org/abs/2608.20373