Skip to content

Comment on Ask HN: Why can't image generation models spell?

Comments

It's the dataset most of the images that were tagged don't have the text shown in the captions. We do a lot of car loras and if you tag the shown text as an example on the numberplate you can prompt/replace it in your prompt without problem.

Newer models like cascade or SD 3 are using multimodal llms to caption images including text. Dall-E was at the forefront because they had access to gpt4-vision before everyone else. You will see that all new models will be able to spell. The problems we see are still mostly because of gigo.

This was my feeling too. Considering a bunch of new images models are coming out at once claiming they can all spell now implies to me it was likely just a training set issue - the caption generators just needed to be told to include any text in the images.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.