AI Image Generators Lose Track of Training Data at Scale
MIT researchers find that large datasets make it impossible to trace generated images back to specific source material, complicating copyright disputes.

A new study from MIT researchers challenges a fundamental assumption in debates over AI-generated art: that you can trace an output back to the training data that influenced it. For models trained on sufficiently large datasets, that connection often doesn't exist at all.
Researchers at MIT's Computer Science and Artificial Intelligence Laboratory identified what they call "attribution decay" — a phenomenon where expanding training data makes individual examples progressively less relevant to any given output. Remove a single image, an entire artist's portfolio, or every photo of a specific person from the training set, and the generated result stays the same.
Zheng Dai, who led the research, frames the finding starkly: if deleting something changes nothing, it can't be responsible for anything. The team tested this across every piece of training data and found the output remained unchanged regardless of what they removed.
Why it matters
This research arrives as artists sue AI companies, legislators draft attribution requirements, and courts weigh whether generated images constitute derivative works. If outputs genuinely can't be traced to specific inputs at scale, existing legal frameworks built on provenance and attribution may not apply. The finding also suggests a path forward for AI companies: architectures that guarantee unattributable outputs could support fair use defenses while protecting individual creators from unauthorized reproduction.
Testing attribution through deletion
Previous attribution methods relied on approximations that estimated influence without actually removing training data. MIT's approach uses a "diffusion ensemble" — multiple smaller models each trained on different data slices. To test what a model would produce without a specific image, researchers simply deactivated the components that had seen it, creating a true counterfactual without retraining.
The team trained 24 ensembles on datasets ranging from 256 to over 160,000 images from public collections including CIFAR-10, CelebA, and ArtBench. The pattern held across all scales: larger training sets produced smaller "counterfactual radii," following an inverse power law. The relationship remained consistent whether measured by pixel differences or semantic meaning.
Stress tests confirmed the finding. Researchers retrained 1,282 separate models the brute-force way at small scale — attribution decay appeared regardless of method. They tested across fixed removal fractions, different training epochs, text-prompted models, and four similarity metrics. The decay survived every variation.
Legal and technical implications
MIT Professor David Gifford, a principal investigator on the project, sees direct relevance to copyright law. If model outputs have no connection to individual training examples, questions arise about fair use, whether outputs qualify as copyrightable novel works, and how to compensate creators when attribution is mathematically impossible.
Cornell Law Professor James Grimmelmann notes the research suggests attribution will fail for sophisticated models, forcing courts and technologists toward other methods for assessing copying versus coincidence.
The study examined diffusion models, now dominant in image generation and scientific applications from protein modeling to drug discovery. Whether the same decay affects large language models at the center of major copyright litigation remains unknown.
The research, supported by Schmidt Futures, was published in Nature Communications and first reported by MIT News.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call