NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Researchers have developed a new framework called NaviDC-OCR for parsing documents. The system addresses two major challenges in document parsing: accurately analyzing the layout of camera-captured documents and dealing with high-resolution scenarios where existing methods often produce redundant or incorrect results. NaviDC-OCR uses deformation-aware learning to improve geometric perception, an adaptive sampling mechanism for complex layouts, and a content-structure decouple
Researchers have developed a new framework called NaviDC-OCR for parsing documents. The system addresses two major challenges in document parsing: accurately analyzing the layout of camera-captured documents and dealing with high-resolution scenarios where existing methods often produce redundant or incorrect results. NaviDC-OCR uses deformation-aware learning to improve geometric perception, an adaptive sampling mechanism for complex layouts, and a content-structure decoupled learning strategy to model formula grammars and table structures. The framework achieves state-of-the-art performance on various benchmarks, including the ICDAR 2026 Sci-ImageMiner Challenge.
---
Why it matters: This matters because accurate document parsing is crucial for many applications, such as data extraction from historical documents or medical records. NaviDC-OCR's ability to handle complex layouts and high-resolution scenarios makes it a significant advancement in this area.
Source: https://arxiv.org/abs/2608.12898
This article was originally published at: https://arxiv.org/abs/2608.12898