Skip to content

Comment on Ask HN: What's the current best way to extract tables from PDFs?

Comments

Field report- the problem is subtle. I wrote code to do this for mine, rather than use CSVs, because the statement is a regulated document, which CSVs are not, and it has balances for validation, which CSVs also lack.

I wound up with a pipeline of pdftotext -> configurable regexes to capture the transactions within their respective sections (banks list credits and debits separately without indicating the sign in the amount field) -> BNF parser to turn transaction lines into data, then checks start balance + transactions = end balance.

PITB but works well.

Over the winter will be standing up a local model to see whether a sophisticated prompt can reliably accomplish the same.

Not going to base any workflow on my transaction data on hosted models.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.