Everybody's everything is being used to train "AI".
That was the reason for twitter going logged-in-only. Multiple companies were trying to download all of twitter at the same time to train their models.
Reddit put walls up for the same reason, though as usual they lied about why they were making reddit worse.
That's just the state of the Internet right now: If you have a large dataset of reasonably quality content, it's going to get slurped up to train models.
Which is vaguely funny, and feels like a gut response from the business side of the house without understanding the real tech.
If you shut your data off, but then sell it to be used to generate an LLM, you have then given your data away again.
In the simplest case, as the weights directly.
In the harder cases, as the use case it's being applied to solve... ex: if you make an excellent gardening help chat bot, I can clone yours by training my chat bot on your chat bot (which is what a ton of the open/free models are doing right now).
I now have the "jpg" version of your dataset. Is it as good? Meh, probably not. Is it good enough? Almost certainly.
So basically my addendum to your comment is really:
That's just the state of the Internet right now: If you have a large dataset of reasonably quality content, it's going to get slurped up to train models.
And then those models are getting slurped up to train open models. There IS NO MOAT!!!
That was the reason for twitter going logged-in-only.
It's probably one of the reasons for why they did it initially (although now they reverted it so viewing individual tweets can be done by guests), but I think it's guaranteed that they wanted to squeeze out more user registrations and create an appearance of growth.
Comments
Everybody's everything is being used to train "AI".
That was the reason for twitter going logged-in-only. Multiple companies were trying to download all of twitter at the same time to train their models.
Reddit put walls up for the same reason, though as usual they lied about why they were making reddit worse.
That's just the state of the Internet right now: If you have a large dataset of reasonably quality content, it's going to get slurped up to train models.
Which is vaguely funny, and feels like a gut response from the business side of the house without understanding the real tech.
If you shut your data off, but then sell it to be used to generate an LLM, you have then given your data away again.
In the simplest case, as the weights directly.
In the harder cases, as the use case it's being applied to solve... ex: if you make an excellent gardening help chat bot, I can clone yours by training my chat bot on your chat bot (which is what a ton of the open/free models are doing right now).
I now have the "jpg" version of your dataset. Is it as good? Meh, probably not. Is it good enough? Almost certainly.
So basically my addendum to your comment is really:
And then those models are getting slurped up to train open models. There IS NO MOAT!!!
So it was just a convenient side effect that it killed off all the good third party apps to push people towards their app.
It's probably one of the reasons for why they did it initially (although now they reverted it so viewing individual tweets can be done by guests), but I think it's guaranteed that they wanted to squeeze out more user registrations and create an appearance of growth.
So perhaps we need to require that commercial models trained on public data be also made public?