Yeah, sorry for confusion. When said Unicode, meant foreign text rather (just) the unescaped symbols, e.g. Greek. At one random Greek textbook[0], zpdf output is (extract | head -15):
Lol, but there's 100 competitors in the PDF text extraction space, some are multi million dollar industries: AWS textract, ABBY PDFreader, PDFBox, I think you may be underestimating the challenge here.
Comments
In my experience with parsing PDFs, speed has never been an issue, it has always been a matter of quality.
I tried a small PDF and got a memory error. It's definitely much faster than MuPDF on that file.
“The fastest PDF extractor is the one that crashes at the beginning of the file” or something.
fixed.
Yeah, sorry for confusion. When said Unicode, meant foreign text rather (just) the unescaped symbols, e.g. Greek. At one random Greek textbook[0], zpdf output is (extract | head -15):
This for entire book. Mutool extracts the text just fine.[0]: https://repository.kallipos.gr/handle/11419/15087
works now!
ΑΛΕΞΑΝΔΡΟΣ ΤΡΙΑΝΤΑΦΥΛΛΙΔΗΣ Καθηγητής Τμήματος Βιολογίας, ΑΠΘ
Nice! Speed wasn't even compromised. Still 5x when benching. Also saw now there's page with tool compiled to wasm. Cool.
thanks! :)
sorry, I haven't yet figured out non-latin with tounicode references.
Lol, but there's 100 competitors in the PDF text extraction space, some are multi million dollar industries: AWS textract, ABBY PDFreader, PDFBox, I think you may be underestimating the challenge here.