AI

Source-Free MT Evaluation Is Not MT Evaluation

Machine translation evaluation methods often rely on reference-based metrics, which can be biased and unfair to systems that preserve the source meaning. Researchers argue that adequacy should be judged with respect to the source text, rather than relying solely on references. They propose reframing quality estimation as a primary approach for evaluating machine translation adequacy, and designing hybrid metrics that prioritize source-hypothesis faithfulness while using refer
Machine translation evaluation methods often rely on reference-based metrics, which can be biased and unfair to systems that preserve the source meaning. Researchers argue that adequacy should be judged with respect to the source text, rather than relying solely on references. They propose reframing quality estimation as a primary approach for evaluating machine translation adequacy, and designing hybrid metrics that prioritize source-hypothesis faithfulness while using references as auxiliary evidence. --- Why it matters: This matters because it highlights the limitations of current evaluation methods in machine translation, which can lead to unfair comparisons between systems. By prioritizing source-grounded adequacy evaluation, researchers can develop more accurate and unbiased metrics for evaluating machine translation quality. Source: https://arxiv.org/abs/2608.20925

This article was originally published at: https://arxiv.org/abs/2608.20925