Skip to content

Comment on Tabula: Extract Tables from PDFs

Comments

I have a bunch of scanned PDFs from an open data request I'm looking forward to trying when I'm home. My own solution with pytesser was pretty effective but required a ton of tweaking.

I don't think it's going to be able to help you if they're scans. From the README:

Tabula only works on text-based PDFs, not scanned documents. If you can click-and-drag to select text in your table in a PDF viewer (even if the output is disorganized trash), then your PDF is text-based and Tabula should work.

Ah thanks, missed that. They gave me half text based and half scans, gotta love the government.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.