AI

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

Researchers have developed a framework to analyze speeches from Sri Lankan parliamentary debates. The debates are in three languages: Sinhala, Tamil, and English, and the text is often embedded within complex PDF layouts with mixed scripts and morphology. A deep learning-based approach extracts text from these documents and uses clustering to identify topics across all three languages without requiring explicit labels or supervision. This method successfully recovered 30 macr
Researchers have developed a framework to analyze speeches from Sri Lankan parliamentary debates. The debates are in three languages: Sinhala, Tamil, and English, and the text is often embedded within complex PDF layouts with mixed scripts and morphology. A deep learning-based approach extracts text from these documents and uses clustering to identify topics across all three languages without requiring explicit labels or supervision. This method successfully recovered 30 macro-topics and showed temporal alignment with significant national events. --- Why it matters: This research matters because it enables the analysis of complex, multilingual datasets that were previously inaccessible. The ability to extract insights from these debates can inform understanding of social and political dynamics in Sri Lanka. Source: https://arxiv.org/abs/2608.20365

This article was originally published at: https://arxiv.org/abs/2608.20365