Skip to content

Comment on Tabula: Extract Tables from PDFsparent

Comments

Hi. Tabula author here.

We use JPedal for rendering pages as images. For parsing, we use Apache PDFBox. In the near future, we plan to render the PDFs client side with Mozilla's PDF.js

It's worth mentioning that PDFBox 2.0 does a great job of rendering PDFs too.

PDFBox 1.8 less-than-great rendering engine forced us to include a separate library for that purpose only.

Moving to PDFBox 2.0 is also on our roadmap. But the text extraction API in 2.0 has changed a lot too, so porting our engine would require quite a bit of effort.

Friendly reminder: we're an MIT-licensed open source project, and we're always open to contributions!

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.