Comment on SmolDocling: An ultra-compact VLM for end-to-end multi-modal document conversionparentComments−nanoxid1yOCR is not the task being solved here, though. This is supposed to help you when dealing with complex layouts where text is not just read left-to-right, top-to-bottom.But I agree that accurate OCR is kind of a prerequisite for adaptation.
Comments
OCR is not the task being solved here, though. This is supposed to help you when dealing with complex layouts where text is not just read left-to-right, top-to-bottom.
But I agree that accurate OCR is kind of a prerequisite for adaptation.