Skip to content

Comment on Efficient high-resolution image synthesis with linear diffusion transformer

Comments

Does this finally solve the class of "6 fingers/hand" problems?

That problem can be fixed through careful fine-tuning, at the cost of losing some generality because the model is punished for drawing bad fingers. This new method outlined in the paper operates in a highly spatially-compressed latent space, but with more channels than previous models, so each latent pixel has 2x the information content than Flux and 8x the content of SDXL. I do wonder whether the high spatial compression means that high resolution features like fingers will be messed up. On the other hand, the higher channel count in the latent space gives the model more detail per pixel to work with… I guess we’ll just have to see.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.