Skip to content

Comment on Digitizing 55,000 pages of civic meetings

Comments

Hrmm, Berkeley resident here. It's clear that the city possesses the original digital copies of the last several decades of these documents, and the fact that their system of record produces them as images is just a weird quirk. Instead of OCRing them, wouldn't it be better to just get the city to fix their system? I'm not the only one who thinks so. Berkeleyside recently wrote about it:

https://www.berkeleyside.org/2022/08/12/new-city-website-lim...

FOIA request for the entire corpus on digital media and upload to the Internet Archive as a collection? The archive OCRs PDFs as part of derive operations when items are uploaded. Could even crowdsource the FOIA request using Muckrock.com.

Ok but the originals were available on the city's website until literally a few weeks ago. I don't think disappearing them was intended, it was just the inevitable result of hiring some consultant clowns to redo the site. It's a very open city, with a full-time staffed archive at the central library who will find whatever you want to read.

I was not aware. Hopefully these docs can be returned to their previously publicly available glory in short order. Putting them in the Internet Archive ensures access in perpetuity.

I’m not too clear on the intricacies of FOIA requests but some data cities have, they charge for. Would FOIA be able to get that as well?

Otherwise I feel this type of request could be denied for being too large/expensive in nature to fulfill.

For instance many counties provide historical tax records for a fee, companies like Zillow pay for this data. Curious if you know more for the purpose of freeing more data.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.