No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation
Researchers have developed a method to create and validate unknown person names for use in evaluating language models' factuality, privacy leakage, bias, and other traits. The Plausible Unknown Names (PUN) protocol combines data from Wikidata with web-enabled screening and controlled search revalidation to ensure the generated names are plausible but not identifiable. A human study found that participants were able to recover person evidence in only 3% of cases using these na
Researchers have developed a method to create and validate unknown person names for use in evaluating language models' factuality, privacy leakage, bias, and other traits. The Plausible Unknown Names (PUN) protocol combines data from Wikidata with web-enabled screening and controlled search revalidation to ensure the generated names are plausible but not identifiable. A human study found that participants were able to recover person evidence in only 3% of cases using these names.
---
Why it matters: This matters because it provides a more reliable way to evaluate language models' performance on sensitive tasks, reducing the risk of biased or inaccurate results due to memorization or retrieval of specific names. It also enables researchers to better understand how models handle unknown or ambiguous inputs.
Source: https://arxiv.org/abs/2608.21206
This article was originally published at: https://arxiv.org/abs/2608.21206