Skip to content

Comment on The Pile: An 800GB dataset of diverse text for language modeling (2020)parent

Comments

If an individual trains a model on their own data to embody their own skills and behaviour, so that they can then sell/rent that model out to work on their behalf

No, because we already do not treat all work as copyrightable. A plumber doesn't get copyright on his piping job. It has to be original enough. So while your own skill might be original enough to warrant copyright, distilling it into a model might not.

An artists work is copyrightable, a writers work is copyrightable, an a personal model could reproduce those and also produce new works in the same style. Also, data can be intellectual property without being copyrightable.

an a personal model could reproduce those and also produce new works in the same style

Yes. So it's like creating a machine that can create art.

Perhaps it shouldn't be copyrightable, but patentable. I think I would be OK with ML models (weights) being patentable rather than copyrightable.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.