"The data we licensed from Shutterstock was critical to the training of DALL-E,” said Sam Altman
So there was legitimate Shutterstock data used to train DALL-E. I don't think I've read this before. The images with Shutterstock copyright watermarks were presumably scraped accidentally from the web.
The contributor fund is also setting a precedent that data used to train a model has a claim for payment.
Shutterstock is launching a “Contributor Fund” that will reimburse creators when the company sells work to train text-to-image AI models. This follows widespread criticism from artists whose output has been scraped from the web without their consent to create these systems. Notably, Shutterstock is also banning the sale of AI-generated art on its site that is not made using its DALL-E integration.
It's also possible that they licensed the images after the fact so they wouldn't get sued, and what Altman is describing _are_ the watermarked images. It's possible Shutterstock got very upset with them and they worked this out quietly, outside of court.
Comments
From https://www.shutterstock.com/press/20435?irclickid=UTywLoXV1...
So there was legitimate Shutterstock data used to train DALL-E. I don't think I've read this before. The images with Shutterstock copyright watermarks were presumably scraped accidentally from the web.
The contributor fund is also setting a precedent that data used to train a model has a claim for payment.
I can't imagine the payout from a contributor fund for your image being used in training would be anything but pennies or less.
"DALL·E 2 is trained on hundreds of millions of captioned images from the internet."
It's also possible that they licensed the images after the fact so they wouldn't get sued, and what Altman is describing _are_ the watermarked images. It's possible Shutterstock got very upset with them and they worked this out quietly, outside of court.