PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference Training
Researchers have proposed a new framework called PrivAct to improve the privacy of large language models in personalized tasks. The existing approaches to protecting sensitive information rely on external interventions that can be brittle and scenario-specific. PrivAct internalizes contextual privacy preservation into the model's generation behavior, embedding privacy preferences into each agent. This approach enhances system-wide contextual integrity while maintaining compar
Researchers have proposed a new framework called PrivAct to improve the privacy of large language models in personalized tasks. The existing approaches to protecting sensitive information rely on external interventions that can be brittle and scenario-specific. PrivAct internalizes contextual privacy preservation into the model's generation behavior, embedding privacy preferences into each agent. This approach enhances system-wide contextual integrity while maintaining comparable helpfulness. Experiments show consistent improvements in contextual privacy preservation, reducing leakage rates by up to 12.32%. The code for PrivAct is available on GitHub.
---
Why it matters: This matters because large language models are increasingly used in tasks that involve sensitive information, and current methods of protecting this data can be flawed. PrivAct's framework provides a more robust way to preserve contextual privacy, which is crucial for maintaining trust in AI systems.
Source: https://arxiv.org/abs/2602.13840
This article was originally published at: https://arxiv.org/abs/2602.13840