Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models
Researchers have proposed corrections to Zipf's and Heaps' laws, which describe the distribution of word frequencies in language. The corrections are based on a model that takes into account the proportion of words that occur only once (hapaxes) and their relationship to text length. Four different models were tested, with one showing the best fit for a sample of 14 English texts. The study suggests that more complex models may be needed to accurately describe larger language
Researchers have proposed corrections to Zipf's and Heaps' laws, which describe the distribution of word frequencies in language. The corrections are based on a model that takes into account the proportion of words that occur only once (hapaxes) and their relationship to text length. Four different models were tested, with one showing the best fit for a sample of 14 English texts. The study suggests that more complex models may be needed to accurately describe larger language corpora.
---
Why it matters: This research matters because it provides new insights into the statistical properties of language and can inform the development of more accurate natural language processing models.
Source: https://arxiv.org/abs/2307.12896
This article was originally published at: https://arxiv.org/abs/2307.12896