Skip to content

Comment on The Pile: An 800GB dataset of diverse text for language modeling (2020)parent

Comments

Stay tuned! I've got a paper I'm writing about a new followup which is a 40x improvement in size (basically every open source debate card... Ever) and a 40x improvement in metadata and duplication detection. The work is all done since late april and I've just been lazy/writer-blocked (ironic in a world of high end LLMs) and haven't gotten the paper finished.

Kinda of sad to have missed NeurIPS dataset track deadline and ACL, but I know that anything close to this in scope is a slam-dunk accept at the argument mining workshop

Would love to see an early version of it!

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.