SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
Researchers propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework for evaluating large language models. Unlike traditional evaluation methods that rely on holistic A/B preference choices, SESSE breaks down the evaluation process into structured sub-questions based on the judge's own error cases. This approach requires no oracle responses, task-specific rubrics, or fine-tuning and achieves competitive results with a fine-tuned specialist model.
Researchers propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework for evaluating large language models. Unlike traditional evaluation methods that rely on holistic A/B preference choices, SESSE breaks down the evaluation process into structured sub-questions based on the judge's own error cases. This approach requires no oracle responses, task-specific rubrics, or fine-tuning and achieves competitive results with a fine-tuned specialist model.
---
Why it matters: SESSE matters because it provides a more nuanced understanding of large language models' strengths and weaknesses by isolating specific quality dimensions that drive evaluation outcomes. This can help researchers identify areas for improvement and develop more effective training strategies.
Source: https://arxiv.org/abs/2608.18303
This article was originally published at: https://arxiv.org/abs/2608.18303