Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
Researchers have evaluated the performance of open-source Optical Character Recognition (OCR) engines and Large Language Models (LLMs) on a complex task in the public sector. They tested these models on extracting information from student applications for an international study program, which is considered a high-risk application under EU AI regulations. The results show that even state-of-the-art open-source models struggle to handle this task reliably, with most configurati
Researchers have evaluated the performance of open-source Optical Character Recognition (OCR) engines and Large Language Models (LLMs) on a complex task in the public sector. They tested these models on extracting information from student applications for an international study program, which is considered a high-risk application under EU AI regulations. The results show that even state-of-the-art open-source models struggle to handle this task reliably, with most configurations scoring below 0.25. Input quality and OCR output preservation are critical factors affecting performance.
---
Why it matters: This research matters because it highlights the challenges of using open-source AI tools in high-risk applications, such as those in the public sector. It shows that even top-performing models may not guarantee reliable results, emphasizing the need for careful evaluation and consideration of input quality.
Source: https://arxiv.org/abs/2608.18289
This article was originally published at: https://arxiv.org/abs/2608.18289