Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
Researchers have created a multilingual parallel corpus with annotations of multi-word expressions. The corpus, called AlphaMWE, includes translations and annotations for six languages: Arabic, Chinese, English, German, Italian, and Polish. The team used machine translation followed by human post-editing to create the corpus, which they believe will be useful for research in areas such as multi-word term lexicography, machine translation, and information extraction. They also
Researchers have created a multilingual parallel corpus with annotations of multi-word expressions. The corpus, called AlphaMWE, includes translations and annotations for six languages: Arabic, Chinese, English, German, Italian, and Polish. The team used machine translation followed by human post-editing to create the corpus, which they believe will be useful for research in areas such as multi-word term lexicography, machine translation, and information extraction. They also identified challenges that machine translation systems face when translating multi-word expressions.
---
Why it matters: This matters because it provides a valuable resource for researchers working on multilingual natural language processing tasks, allowing them to train and evaluate models on a large-scale annotated dataset.
Source: https://arxiv.org/abs/2011.03783
This article was originally published at: https://arxiv.org/abs/2011.03783