The Authenticity Gap in Human Evaluation
A recent study on natural language generation (NLG) evaluation has found that the standard method of collecting human ratings may not accurately reflect human preferences. The researchers analyzed the assumptions behind this approach and discovered that it can be flawed when using Likert scales, which can reverse the direction of true preference in certain cases. To address this issue, the authors propose a new protocol called system-level probabilistic assessment (SPA) for e
A recent study on natural language generation (NLG) evaluation has found that the standard method of collecting human ratings may not accurately reflect human preferences. The researchers analyzed the assumptions behind this approach and discovered that it can be flawed when using Likert scales, which can reverse the direction of true preference in certain cases. To address this issue, the authors propose a new protocol called system-level probabilistic assessment (SPA) for evaluating open-ended tasks like story generation.
---
Why it matters: This matters to researchers in AI because it highlights the limitations of current human evaluation methods and suggests that existing results may not be reliable. The proposed SPA protocol could improve the accuracy of NLG evaluations, but its applicability is limited to specific tasks.
Source: https://arxiv.org/abs/2205.11930
This article was originally published at: https://arxiv.org/abs/2205.11930