AI

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Researchers have investigated the problem of specifying a reward function for large language models (LLMs) to unlearn target-specific knowledge. They conducted experiments in a controlled setting using four different reward designs and found that optimization success does not necessarily mean behavioral unlearning. The study highlights discrepancies between various metrics, including RWKU forget scores, held-out completion audits, and training dynamics. These findings suggest
Researchers have investigated the problem of specifying a reward function for large language models (LLMs) to unlearn target-specific knowledge. They conducted experiments in a controlled setting using four different reward designs and found that optimization success does not necessarily mean behavioral unlearning. The study highlights discrepancies between various metrics, including RWKU forget scores, held-out completion audits, and training dynamics. These findings suggest that current benchmark probes may miss endpoint changes and rewards can inadvertently select broad-topic answering with low semantic leakage. --- Why it matters: This research matters to engineers working on large language models because it reveals the limitations of current reward functions and benchmarking methods. The study's findings highlight the need for more accurate and reliable evaluation metrics, which is crucial for advancing the field of LLM unlearning. Source: https://arxiv.org/abs/2608.17804

This article was originally published at: https://arxiv.org/abs/2608.17804