Skip to content

Comment on Meet PastPages.org, the news homepage archive. (And help keep it alive)

Comments

I cre­ated this site be­cause I think it ought to ex­ist. The shift­ing homepages of ma­jor me­dia sites should be saved so they can be stud­ied. Done right, I be­lieve Pas­t­Pages could serve as a re­source for schol­ars seek­ing to study cov­er­age of news events, like the up­com­ing U.S. pres­id­en­tial elec­tion.

Collecting this data cost money. So I've set up a Kickstarter drive to raise funds. If you'd like to help keep PastPages alive, please considering giving.

http://www.kickstarter.com/projects/651552740/keep-pastpages...

What stack are you using for the web page capture? It's a perfect crisp capture. I've tried before and never got such good programatic results.

I'm using Selenium's Firefox driver from inside a Django app. There is Python binding that's slick once you figure out a couple timeout related workarounds that are necessary. Their forums helped me over that hurdle.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.