Fast Company
Follow
Does generative AI actually copy artists? Researchers say it’s up for debate
A new MIT study suggests generative AI might not directly copy specific artworks. Researchers investigated diffusion models, commonly used for image generation, to see if their outputs could be traced to individual training data pieces. They discovered that larger training datasets make it significantly harder to link generated images to any single source. Removing a specific piece of training data often had minimal impact on the model's output, a phenomenon called "attribution decay." This means AI can create images resembling an artist's style without a provable causal link to that artist's specific contribution. The researchers theorize this occurs because large datasets contain visual redundancy, with many images sharing similar features. They developed a method called "ablation" to test cause and effect by training model components separately. This allowed them to isolate and remove the influence of specific data without full retraining. Their experiments showed that assuming a source based on visual similarity alone is often inaccurate. The "unattributability" effect is more pronounced with extensive datasets, with significant impact seen at scales of tens of thousands of images. Using Andy Warhol's art as an example, removing his silkscreens from a large dataset would not drastically alter the generated output if similar styles were present elsewhere. The study's findings could weaken arguments that AI has directly copied an artist's specific work, as proving a definitive source becomes challenging. While offering a potential defense for AI companies in court regarding specific outputs, it does not address the legality of using copyrighted material for training.