Skip to content

Comment on How to generate realistic people in Stable Diffusion

Comments

Try using the "I Can't Believe It's Not Photography" model. Instead of trying to micro-manage the details, use strong emotional terms. I've had good results with prompts along the lines of "Aroused angry feral Asian woman wearing a crop top riding a motorcycle fast in monsoon rain."[1]

[1] https://i.ibb.co/3zHGyrR/feral34.png

RealvisXL is much better, probably the best model for photorealism. Example with same prompt: https://gencdn.aieasypic.com/original/ed7c2742-4a76-4e22-9bb...

It’s currently the best open weights model for prompt adherence too: https://imgsys.org/

It’s currently the best open weights model for prompt adherence too: https://imgsys.org/

This doesn't specifically measure prompt following but how good it is overall.

Is it also based on SD?

Pro tip, if you see a model with XL on the end, it's for SDXL, which is a stable diffusion model

That looks incredibly fake, to be fair.

The rain is heavy enough to be coming off her body in sheets, but not heavy enough to have plastered her hair onto her head yet.

The bike appears to have 2 front brake levers, only one fork, and she is holding the left handlebar the wrong side of the switches.

And you would say this is realistic?

If a person had drawn or painted it by hand, or put it together in Blender, I would say they were aiming for a realistic style, and they'd done an impressive job.

Sure the motorbike handlebar is messed up, as is one of the hands. And the torso doesn't quite match up with where the thigh is. And the reflections on the arms, the face and the cleavage all look differently lit. And the hair isn't behaving like wet hair. And the face has an uncanny-valley airbrushed look about it.

But those are all easily overlooked. Chuck this into a Facebook news feed and I think 70% of the general public would believe it was a photograph.

While I think the number is much smaller than 70%, you make a good point: Relative to a human artist it is pretty realistic, especially relative to a painter who doesn't use Photoshop or Blender.

It has life and energy. If you put static descriptive phrases into a Stable Diffusion system, you get a static, boring scene. Stable Diffusion can do better than that.

Some people do have six fingers.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.