PersonalBench: Measuring the Authorship Gap in LLM Personalization
Researchers have introduced PersonalBench, a benchmark for evaluating personalized text generation in large language models (LLMs). The benchmark assesses whether an LLM can mimic the writing style of a specific individual. To do this, it uses three methods: a trained authorship verification model called LUAR, an LLM that judges the output, and automated stylometrics. The results show that while personalization methods can produce text that resembles a target author's writing
Researchers have introduced PersonalBench, a benchmark for evaluating personalized text generation in large language models (LLMs). The benchmark assesses whether an LLM can mimic the writing style of a specific individual. To do this, it uses three methods: a trained authorship verification model called LUAR, an LLM that judges the output, and automated stylometrics. The results show that while personalization methods can produce text that resembles a target author's writing, they still fall short of human-like writing. In fact, the LLM's own 'authorship fingerprint' dominates its generated text, making it more distinct from any human author than random humans are from each other. The study suggests that current personalization methods do not bridge the gap between machine and human authorship.
---
Why it matters: This research matters to AI engineers because it highlights the limitations of current personalized text generation techniques. It shows that even with advanced models, there is still a significant gap between machine-generated text and human writing. This has implications for applications such as content creation, where the ability to mimic human-like writing is crucial.
Source: https://arxiv.org/abs/2608.19746
This article was originally published at: https://arxiv.org/abs/2608.19746