Skip to content

Comment on Efficient high-resolution image synthesis with linear diffusion transformerparent

Comments

That transfers computer time to user time. It's great when you want variations, less so when you want precision and consistency. Picking the best image tires the brain quite quickly, you have to take into account the at a glance quality without it overriding the detail quality.

I'd be curious to see how a vision model would go if it were finetuned to select the best image match to a given criteria.

It's possible that you could do O1 style training to build a final stage auto-cherrypicker.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.