Skip to content

Comment on Show HN: Simply explain 20k concepts using GPT

Comments

The amount of content thats going to be generated in the next few years is going to absolutely drown anything humanity has created thus far. I would not be surprised to see a wholesale return to analog/"old" knowledge once the vast majority on the network becomes unreliable/generated.

Yes, but this also assumes that a lot of the content created by humanity has been worthwhile. A lot of it is pure noise. Youtube is a really great example of this, in between the "hidden gems" and "popular high quality" works there's a whole bunch of noise that might as well have been generated by AI. Actually I wouldn't be surprised if they were at least partially AI generated, they just make some cheap voice actor read their generated script for them.

I think you're going to see a very rapid pendulum swinging here. With an absurd amount of content being generated and then flooding the various platforms, and in turn platforms are going to try and combat this by creating more centralized sources of information. The return to analog knowledge seems a bit far-fetched. I highly doubt that would be an outcome, if only because convenience trumps it. Look at librarians, you can talk to one and get much better direction of information than asking google, but few people do that. I can't see that changing.

Yes, but this also assumes that a lot of the content created by humanity has been worthwhile. A lot of it is pure noise.

And all of those models have been indiscriminately trained on the sum total of that pure noise. What you are getting from these models is a cleaned-up, grammatically-perfect, auto-editorialized synthesis of that pure noise.

Right now, the ability for people to put out shit content is limited by human timescales.

AI algorithms can generate 'noise' at a much faster rate than people can, so it will be even harder to find the hidden gems.

And what happens when the models themselves are trained primarily on the data they produced?

They will spit out crappy, copycat material which might be fake. Like 90% of the material created by humans. And then the people with something valuable to say will have something original and valuable. And we'll be back to having a search/curation/review problem on our hands.

I don't learn from 90% of material created by humans.

I learn from the 1% of material that is created by humans who are, to some degree, experts in their field.

Sometimes it's also wrong, but it's not because they were just lazily regurgitating rando 'net posts.

We already pretrain them on random internet crap then finetune on supervised data

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.