Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale
Researchers from Harvard University have developed an open-source pipeline called Enriched Text that can handle large collections of digitized books. The pipeline preserves metadata and allows users to customize the output based on their needs. It includes features such as detecting duplicate paragraphs, identifying per-paragraph language, and computing bits-per-byte scores. This approach is a response to existing pipelines that often discard meaningful metadata or restrict b
Researchers from Harvard University have developed an open-source pipeline called Enriched Text that can handle large collections of digitized books. The pipeline preserves metadata and allows users to customize the output based on their needs. It includes features such as detecting duplicate paragraphs, identifying per-paragraph language, and computing bits-per-byte scores. This approach is a response to existing pipelines that often discard meaningful metadata or restrict by language. The Enriched Text pipeline can be applied to over 250 languages and is designed to make large collections of digitized books easier for both humans and machines to parse.
---
Why it matters: This matters because it provides a customizable solution for handling large collections of digitized books, which can be useful for researchers and developers working with multilingual texts. The pipeline's ability to preserve metadata and allow users to customize the output based on their needs can also help reduce duplication of effort in processing and analysis.
Source: https://arxiv.org/abs/2608.19026
This article was originally published at: https://arxiv.org/abs/2608.19026