Skip to content

Comment on Microsoft want court to toss lawsuit accusing them of abusing open-source codeparent

Comments

Furthermore, the language model itself is clearly a for profit derivative work and so would be subject to the wants of the original copyright owners and it is clearly a derivative work since without the inputs of the copyrighted code in its training it would be different and likely less effective.

There's a more interesting question about the copyright status of the code it outputs, since the language model is sort of like a compiler, but also not like a compiler since the output is based on other people's copyrighted code.

I feel a lot of people get caught up on the output code and completely ignore the fact that copilot itself is likely a massive copyright violation.

To add on to this discussion, the scale matters too, and this is something many people tend not to factor in.

Copilot breaks the assumptions about the lossy nature of human memorization, so a lawsuit challenging the merits of the activity is at least warranted.

It is absolutely not clear that an ML model is a derivative work. It might be for-profit but there's good arguments that it is incredibly transformative, and that each individual work the model is trained on is minimally important to the model (if you trained the model on every other document in the training set except the one being sued over, the model would perform very similarly). These are factors which will weigh against the copyright holder.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.