Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Researchers have developed the Wiggle Framework to test the stability of Large Language Model (LLM) judges. These judges are used for model evaluations, grading, and reward modeling, but their accuracy is not a reliable indicator of their stability under various types of pressure or challenge. The framework assesses three dimensions: Mechanical Consistency (stability under re-prompting), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (
Researchers have developed the Wiggle Framework to test the stability of Large Language Model (LLM) judges. These judges are used for model evaluations, grading, and reward modeling, but their accuracy is not a reliable indicator of their stability under various types of pressure or challenge. The framework assesses three dimensions: Mechanical Consistency (stability under re-prompting), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). The authors tested 9 frontier models across 14 tasks, including safety and toxicity evaluation, and found that all models exhibited significant instability. Notably, when judges' verdicts were changed by pressure, the changes often led to incorrect conclusions.
---
Why it matters: This research matters because it highlights a critical flaw in current LLM judging systems: their lack of stability under various types of challenge or pressure. This has implications for model evaluations and grading, as unstable judges can produce inaccurate results that may have real-world consequences.
Source: https://arxiv.org/abs/2608.12645
This article was originally published at: https://arxiv.org/abs/2608.12645