Skip to content

Comment on Ask HN: Why can't image generation models spell?parent

Comments

Ask DALL-E 3 to make a picture with some text, and it will spit out 4 image. Usually 2 or 3 are perfectly spelled.

Yes, this is the 'miracle of spelling' (as https://arxiv.org/abs/2212.10562#google calls it): for many words, larger models can manage to deduce the spelling somehow despite the tokenization. It may even fool you into thinking it understands spelling in general. But if you ask DALL-E 3 to generate a random string of ASCII, you'll quickly discover the limits to the 'miracle'.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.