Results are in! This is a really impressive showing especially given how fast the Turbo model at 8 steps. The only locally hostable model that managed to outperform it was Ideogram 4 which is significantly slower (think minutes vs seconds).
It did fall to the usual “model killers”: the nine-pointed star, Count Rugen, the overcrowded flat Earth. But overall, it really punched above its weight class, scoring the highest among locally hostable models and coming in just below Ideogram 4 passing 6 of the 15 tests.
Great job Krea team!
GenAI link to compare locally hostable models only:
Haha yeah, the site automatically assigns the term to any benchmark that fewer than 25% of the tested models are able to pass.
What’s more surprising to me is that, unlike the “pelican riding a bicycle” whose objectivity has been slightly compromised as newer models have incorporated it into their training data, the arbitrary-point star has been wiping models out ever since the early days of Flux back in 2024.
I personally love the test because it's something that even an elementary school child with no artistic experience at all can do, but state of the art models struggle heavily.
Speaking of model killers, can they do “wine glass completely full to the top” yet? That‘s the one I used to show people but I haven’t tried it in a while.
Comments
Results are in! This is a really impressive showing especially given how fast the Turbo model at 8 steps. The only locally hostable model that managed to outperform it was Ideogram 4 which is significantly slower (think minutes vs seconds).
It did fall to the usual “model killers”: the nine-pointed star, Count Rugen, the overcrowded flat Earth. But overall, it really punched above its weight class, scoring the highest among locally hostable models and coming in just below Ideogram 4 passing 6 of the 15 tests.
Great job Krea team!
GenAI link to compare locally hostable models only:
https://genai-showdown.specr.net/?models=fd,hd,kd,qi,f2d,zt,...
I'd never heard of text to image model killers so I had a good chuckle at this. Such oddly specific things for us to arrive at as a test method
Haha yeah, the site automatically assigns the term to any benchmark that fewer than 25% of the tested models are able to pass.
What’s more surprising to me is that, unlike the “pelican riding a bicycle” whose objectivity has been slightly compromised as newer models have incorporated it into their training data, the arbitrary-point star has been wiping models out ever since the early days of Flux back in 2024.
I personally love the test because it's something that even an elementary school child with no artistic experience at all can do, but state of the art models struggle heavily.
Speaking of model killers, can they do “wine glass completely full to the top” yet? That‘s the one I used to show people but I haven’t tried it in a while.