Assessing Quality of Experience in Natural Language Generation of German Text
Researchers have created a new dataset called TextQ-German to evaluate the quality of text generated by natural language processing systems. The dataset includes human ratings and annotations for automatic metrics to assess the perceived quality of German text in tasks like summarization and machine translation. A hybrid approach using transformer-based models and linguistic features was found to perform best, outpacing pure transformer models. This work aims to improve the e
Researchers have created a new dataset called TextQ-German to evaluate the quality of text generated by natural language processing systems. The dataset includes human ratings and annotations for automatic metrics to assess the perceived quality of German text in tasks like summarization and machine translation. A hybrid approach using transformer-based models and linguistic features was found to perform best, outpacing pure transformer models. This work aims to improve the evaluation of natural language generation systems by providing a publicly accessible resource.
---
Why it matters: This matters because current automatic metrics for evaluating text quality are often inadequate, leading to poorly performing NLG systems. By developing better evaluation methods and datasets like TextQ-German, researchers can create more effective NLG systems that align with human expectations.
Source: https://arxiv.org/abs/2608.18888
This article was originally published at: https://arxiv.org/abs/2608.18888