Skip to content

Comment on Frog: OCR Tool for Linuxparent

Comments

The only domain where Tesseract is competitive is for perfect "black text on white paper", it gives pretty poor performance when dealing with colored, distorted text, or even strong page structure effects (tables, etc.).

I wouldn't be surprised if their data set is bigger than the stock tesseract, but part of the OCR process is to preprocess the images.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.