Skip to content

Comment on Tabula: Extract Tables from PDFsparent

Comments

I usually use "pdftotext -layout" and write python or perl code to handle the table extraction.

If I need more detailed formatting information, I use "pdftohtml -xml -fullfontname" and process the resulting xml.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.