Skip to content

Comment on The Pile: An 800GB dataset of diverse text for language modeling (2020)parent

Comments

otherwise we're just going to see everything good disappear behind trade secret

We will see that anyway. All the code I work on commercially is copyrighted and yet a trade secret. Existence of copyright (with the exception of copyleft, but that's subversion) didn't help software to be open sourced.

IMHO allowing models to be copyrighted is basically 18th century enclosures again.

All the code I work on commercially is copyrighted

What do you mean by that? Do you continuously copyright the changes?

Yes, more or less. I am not really sure why we legally do that, I believe it's just another protection in case someone actually copies the code.

I mean everything is copyrighted anyway. It's harder to give up copyright than to keep it, so the main point is despite copyright protections, most companies do not publish their code publicly at all. For code that has to be pushed to clients, most companies even take efforts to obfuscate it.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.