OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
Researchers have developed a system called OriginBlame that tracks the origin of data in AI training datasets at both record and token levels. This allows for more precise removal of sensitive information when requested by contributors, reducing the risk of over-deletion. The system was tested on Wikipedia pages and found to reduce dataset-level over-deletion while adding some overhead to processing time.
Researchers have developed a system called OriginBlame that tracks the origin of data in AI training datasets at both record and token levels. This allows for more precise removal of sensitive information when requested by contributors, reducing the risk of over-deletion. The system was tested on Wikipedia pages and found to reduce dataset-level over-deletion while adding some overhead to processing time.
---
Why it matters: This matters because it addresses a practical challenge in AI development: how to remove sensitive data from training datasets without affecting model performance. By providing more precise forget sets, OriginBlame can help improve the reliability of machine unlearning and reduce the risk of unintended consequences.
Source: https://arxiv.org/abs/2607.13037
This article was originally published at: https://arxiv.org/abs/2607.13037