the question is incorrect. they -can- spell. sometimes very long phrases. five letter words should have no issue. have you tried using ideogram? or even just dalle3 prompted well? https://twitter.com/swyx/status/1765091085943218571
in other words.. what have you actually tried? be specific.
Exactly. The premise was true years ago, definitely. But today it's not hard to get correct spelling out of the top models. Ask DALL-E 3 to make a picture with some text, and it will spit out 4 image. Usually 2 or 3 are perfectly spelled. Lesser or older diffusion models (whatever OP is using) sometimes mix cyrillic and latin letters, or invent plausible-looking letters that don't exist in any language. But think about how they work - they are trained to turn pure pixel noise into a more plausible array of pixels for the prompt. It's pretty close to plausible text - misspelling is a nitpick for what it's getting right. Technology progresses over time.
Ask DALL-E 3 to make a picture with some text, and it will spit out 4 image. Usually 2 or 3 are perfectly spelled.
Yes, this is the 'miracle of spelling' (as https://arxiv.org/abs/2212.10562#google calls it): for many words, larger models can manage to deduce the spelling somehow despite the tokenization. It may even fool you into thinking it understands spelling in general. But if you ask DALL-E 3 to generate a random string of ASCII, you'll quickly discover the limits to the 'miracle'.
Comments
the question is incorrect. they -can- spell. sometimes very long phrases. five letter words should have no issue. have you tried using ideogram? or even just dalle3 prompted well? https://twitter.com/swyx/status/1765091085943218571
in other words.. what have you actually tried? be specific.
Exactly. The premise was true years ago, definitely. But today it's not hard to get correct spelling out of the top models. Ask DALL-E 3 to make a picture with some text, and it will spit out 4 image. Usually 2 or 3 are perfectly spelled. Lesser or older diffusion models (whatever OP is using) sometimes mix cyrillic and latin letters, or invent plausible-looking letters that don't exist in any language. But think about how they work - they are trained to turn pure pixel noise into a more plausible array of pixels for the prompt. It's pretty close to plausible text - misspelling is a nitpick for what it's getting right. Technology progresses over time.
Yes, this is the 'miracle of spelling' (as https://arxiv.org/abs/2212.10562#google calls it): for many words, larger models can manage to deduce the spelling somehow despite the tokenization. It may even fool you into thinking it understands spelling in general. But if you ask DALL-E 3 to generate a random string of ASCII, you'll quickly discover the limits to the 'miracle'.
I tried gpt4, Gemini, stable diffusion. I had sorta assumed DALLE wasn’t the hotness anymore, and had never heard of Ideogram- I’ll try both!