But going back in the '70s with a bunch of 4090s will not help them much, as they wouldn't have enough data to do very meaningful things.
Not at first. But the corpus would have been built over time. Project Gutenberg was founded in 1971. Imagine a world in which the Copyright Act of 1976 requires an electronic copy of every book to be deposited with the Library of Congress.
The best single work of fiction ever created about LLMs' capabilities (and, perhaps, dangers) is Colossus by Jones. Although I think the film is even better than the book, only the latter mentions how, despite being created specifically for US national defense, Colossus is also fed unrelated data including Shakespeare's sonnets, because its creators do not know if it could be important.
"Imagine a world in which the Copyright Act of 1976 requires an electronic copy of every book to be deposited with the Library of Congress."
I fail to imagine such a world. It's 1976, there's very few computers around (at least for the general public and smaller companies), and a lot of typesetting is still not based on computers. Those who are using "digital" typesetting (not Aldus, but interacting with a photo-typesetter via a terminal) are probably storing data in proprietary formats on floppy drives or tapes, and presumably not as a neat .txt file. Who would even consider such a law?
The huge data availability which enabled GPT-2, GPT-3, ChatGPT, etc. is not because of project gutenberg, but because of the capillary usage of Internet and tons of user-produced content you can scrape and reuse. Sure, we can imagine a past in which all these conditions are in place, but then we're just moving all the tech of 2010's and 2020's 50 years behind. And if we do that, I agree we can have current AI approaches and flared trousers go hand in hand :)
Comments
Not at first. But the corpus would have been built over time. Project Gutenberg was founded in 1971. Imagine a world in which the Copyright Act of 1976 requires an electronic copy of every book to be deposited with the Library of Congress.
The best single work of fiction ever created about LLMs' capabilities (and, perhaps, dangers) is Colossus by Jones. Although I think the film is even better than the book, only the latter mentions how, despite being created specifically for US national defense, Colossus is also fed unrelated data including Shakespeare's sonnets, because its creators do not know if it could be important.
"Imagine a world in which the Copyright Act of 1976 requires an electronic copy of every book to be deposited with the Library of Congress."
I fail to imagine such a world. It's 1976, there's very few computers around (at least for the general public and smaller companies), and a lot of typesetting is still not based on computers. Those who are using "digital" typesetting (not Aldus, but interacting with a photo-typesetter via a terminal) are probably storing data in proprietary formats on floppy drives or tapes, and presumably not as a neat .txt file. Who would even consider such a law?
The huge data availability which enabled GPT-2, GPT-3, ChatGPT, etc. is not because of project gutenberg, but because of the capillary usage of Internet and tons of user-produced content you can scrape and reuse. Sure, we can imagine a past in which all these conditions are in place, but then we're just moving all the tech of 2010's and 2020's 50 years behind. And if we do that, I agree we can have current AI approaches and flared trousers go hand in hand :)