Which is vaguely funny, and feels like a gut response from the business side of the house without understanding the real tech.
If you shut your data off, but then sell it to be used to generate an LLM, you have then given your data away again.
In the simplest case, as the weights directly.
In the harder cases, as the use case it's being applied to solve... ex: if you make an excellent gardening help chat bot, I can clone yours by training my chat bot on your chat bot (which is what a ton of the open/free models are doing right now).
I now have the "jpg" version of your dataset. Is it as good? Meh, probably not. Is it good enough? Almost certainly.
So basically my addendum to your comment is really:
That's just the state of the Internet right now: If you have a large dataset of reasonably quality content, it's going to get slurped up to train models.
And then those models are getting slurped up to train open models. There IS NO MOAT!!!
Comments
Which is vaguely funny, and feels like a gut response from the business side of the house without understanding the real tech.
If you shut your data off, but then sell it to be used to generate an LLM, you have then given your data away again.
In the simplest case, as the weights directly.
In the harder cases, as the use case it's being applied to solve... ex: if you make an excellent gardening help chat bot, I can clone yours by training my chat bot on your chat bot (which is what a ton of the open/free models are doing right now).
I now have the "jpg" version of your dataset. Is it as good? Meh, probably not. Is it good enough? Almost certainly.
So basically my addendum to your comment is really:
And then those models are getting slurped up to train open models. There IS NO MOAT!!!