News update
  • Oil Shock Sends Asian Markets Lower as Middle East Conflict Threatens Global Shipping     |     
  • Trump Links Iran War to US Midterms, Backs United Ireland     |     
  • PM Tarique Rahman Seeks Public Cooperation to Tackle Nationwide Dengue Outbreak     |     
  • Jatiyatabadi Mahila Dal Vows to Lead Fight Against Rape, Violence Against Women and Children     |     
  • 'Honorable' Dropped: PM Starts Implementing Decision from His Own Office     |     
Bangolok Desk News 2026-08-20, 4:15pm

MIT study finds AI image outputs often untraceable to individual training images

screenshot_20260820-161030_chrome-1b2251964bf532f66460cc58e82e0f6a1787220900.jpg

Image: Collected


A new study from the Massachusetts Institute of Technology suggests that as image-generating AI models are trained on ever-larger datasets, the influence of any single training image can disappear — a finding that complicates efforts to link AI outputs to specific copyrighted works.

Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) describe the effect as “attribution decay.” Published Tuesday in Nature Communications, the paper reports that in many cases removing one image — or even an entire artist’s body of work — from a model’s training set produced no measurable change in the images the model generated.

The team, led by Zheng Dai and David Gifford, built a testing framework they call a “diffusion ensemble” to ask a simple but costly question: what would a model produce if it had never seen a particular image? Instead of retraining models from scratch for every hypothetical removal, the diffusion ensemble trains multiple overlapping components that can be selectively switched off, letting researchers observe the model’s behavior with and without specific data.

“We designed the system so we could take a piece of data away and see if the output changed,” Dai said in MIT’s announcement. “If the output doesn’t change, then that piece of data didn’t affect the output. So it doesn’t make much sense to attribute the output to that piece of data.”

Testing 24 diffusion ensembles on datasets ranging from 256 to more than 160,000 images, the researchers found a clear pattern: the larger the training set, the less any individual image mattered. In extreme cases, the team reported that removing every photograph of a particular person or an entire artist’s corpus produced no discernible effect on the model’s outputs.

Legal questions arise

The findings arrive amid a wave of copyright litigation targeting AI-image generators. Media giants including Disney, NBCUniversal and DreamWorks have joined suits alleging that companies trained generative models on copyrighted images without permission. High-profile lawsuits include The New York Times’ case against OpenAI and Microsoft, and a 2023 suit against Midjourney and others alleging that scraped images were used to reproduce artists’ styles.

Many of those legal strategies rest on an implicit attribution: that a generated image can be traced back to particular training examples. If attribution often “vanishes” as models scale, courts and litigants may need new methods to establish copying, the study’s authors say.

“We might have to rethink what intellectual property means,” Dai told Semafor. “You can’t just assume it, and the attribution link sort of vanishes.”

Not a blanket defense

Researchers and legal experts caution that the study does not give AI companies a free pass. The MIT team stressed that their work focused on diffusion models — a widely used class of image generators — and does not rule out memorization or direct copying of training images in other circumstances. Past studies have shown that models can memorize and reproduce specific high-frequency or distinctive training examples, particularly when datasets include many near-duplicates or the model is overparameterized.

James Grimmelmann, a law professor at Cornell, said the paper “provides reason to think that attribution will fail for interesting models. Instead, technologists and courts will need to resort to other methods for assessing copying.” Those methods might include forensic image analysis, logs of training inputs, model watermarking, or stronger regulatory disclosure requirements for training datasets.

Implications for artists and policymakers

For artists and photographers pursuing litigation, the study complicates a familiar narrative: that a model’s output is a traceable mosaic of the training set. If attribution is unreliable, plaintiffs could find it harder to prove that a model copied their work, even when generated images bear strong stylistic resemblance.

Policymakers drafting rules for AI training and copyright may also need to adjust. Current proposals range from mandatory opt-in licensing schemes and dataset transparency requirements to new remuneration systems for creators. The MIT results suggest that simple notions of “copying” tied to individual training files may not map neatly onto how modern models actually learn.

Open Questions

The study raises several open technical and legal questions. The research examined diffusion models specifically; whether similar attribution decay occurs in large language models or multimodal systems remains uncertain. The diffusion ensemble approach also relies on a particular construction of model components, and real-world commercial models may behave differently.

The MIT team’s results point to a middle ground: scaling models can make them less dependent on any single example, but that does not eliminate the possibility of copying or harm. As courts and lawmakers weigh the rights of creators against the social benefits of generative AI, new standards for evidence and new technical safeguards will likely become central to disputes over ownership and accountability.