OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
Introduction
FAQ
What is OriginBlame?
A record- and token-level data provenance system that precisely identifies training records belonging to a specific author.
How does OriginBlame solve data deletion?
It resolves revocation requests into precise forget sets via deterministic queries, avoiding dataset-level over-deletion.
What is the impact on model performance?
It improves unlearning by 42% on a 1.7B model with minimal overhead.
Can OriginBlame be used in MENA?
Yes, especially for organizations handling sensitive data and needing compliance with privacy regulations like GDPR.
Source: arXiv cs.AI
AI-assisted content, human-reviewed.