Skip to content

Why the Library of Congress cares about archiving our tweets

arstechnica.com
5 pointsthinkzig5 comments
On HN

Comments

For those interested, I asked around and it looks like the entire Twitter collection is only going to be 5TB worth.

I calculate the tweets data must be around 3.5TB so that figure seems about right (if you throw in user and times tamp data)

I hope they do elect to store the actual URLs when URL shorteners are involved. Really, though it would eat a hole in the URL shorteners' business, I really think sites like Twitter should be dynamically 'fixing' those links before inserting them into the database.

I've been curious. Does this include the tweets of "private" accounts?

The article mentioned that only public accounts would be included.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.