Skip to content

Comment on SmolDocling: An ultra-compact VLM for end-to-end multi-modal document conversion

Comments

After many posts on my feed, I decided to give it a spin.

The good: - Open source.

- Can run locally (Apple Silicon) at a fair speed.

- Image detection is good.

The bad:

- Not detecting tables.

- Text in a perfectly clean PDF (resume) is not detected.

I know its in preview, small and open source which is great, but its far from being usable.

-

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.