AI

When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era

Researchers have created a benchmark called AVShift to evaluate authorship verification under various distribution shifts. These shifts include changes in genre, time, and the use of AI-assisted writing. The benchmark consists of over 150,000 text pairs spanning three genres and 21 years, allowing for controlled evaluation of these factors. Experiments show that fine-tuned large language models (LLMs) generalize best across genres and benefit from diverse training data. Howev
Researchers have created a benchmark called AVShift to evaluate authorship verification under various distribution shifts. These shifts include changes in genre, time, and the use of AI-assisted writing. The benchmark consists of over 150,000 text pairs spanning three genres and 21 years, allowing for controlled evaluation of these factors. Experiments show that fine-tuned large language models (LLMs) generalize best across genres and benefit from diverse training data. However, temporal drift has a significant impact on authorship verification, with performance degrading as the time gap between documents increases. The study also found no evidence of a measurable AI-era distribution shift within AVShift. --- Why it matters: This research matters to engineers and researchers in AI because it highlights the challenges of authorship verification under real-world conditions. Understanding how writing styles change over time and across genres can help improve the robustness of models used for tasks like plagiarism detection and text classification. Source: https://arxiv.org/abs/2608.17979

This article was originally published at: https://arxiv.org/abs/2608.17979