Maybe both are ok. Maybe neither. Maybe one. The point is this does not follow
If training on copyrighted data without authors consent is ok then distilling is ok as well.
Because building a model from distillation and building a model from raw data are not the same. You have to evaluate them independently. And legally its different as well. IP vs ToS (civil).
Comments
So what?
If training on copyrighted data without authors consent is ok then distilling is ok as well.
Maybe both are ok. Maybe neither. Maybe one. The point is this does not follow
Because building a model from distillation and building a model from raw data are not the same. You have to evaluate them independently. And legally its different as well. IP vs ToS (civil).
No, it follows because the AI labs can at best argue that some copyright over the model output was broken.
They can hardly successfully sue a rival for breaking ToS.