The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
Researchers at Netflix have developed a lifecycle framework for Large Language Model (LLM) judges used in recommendation explanations. The framework consists of four phases: building and training the judges, deploying them in production, and continuously monitoring their performance. The LLM judges evaluate user-facing explanations and are refined through a process called Reasoning-Aligned Rubric Tuning (RART). A five-week A/B test showed that judge-aligned explanations incre
Researchers at Netflix have developed a lifecycle framework for Large Language Model (LLM) judges used in recommendation explanations. The framework consists of four phases: building and training the judges, deploying them in production, and continuously monitoring their performance. The LLM judges evaluate user-facing explanations and are refined through a process called Reasoning-Aligned Rubric Tuning (RART). A five-week A/B test showed that judge-aligned explanations increased successful browse-to-play sessions and shifted member viewing toward novel content.
---
Why it matters: This work matters to AI researchers because it highlights the need for a more dynamic approach to evaluating LLMs, which are increasingly used in production systems. The lifecycle framework provides a structured way to build, train, deploy, and maintain LLM judges, addressing technical and operational challenges that arise during each phase.
Source: https://arxiv.org/abs/2608.18300
This article was originally published at: https://arxiv.org/abs/2608.18300