Skip to content

Comment on How to generate realistic people in Stable Diffusion

Comments

I might be going around in the wrong social circles, but none of the people I know look anything like the realistic people in these images. Are these models even able to generate pictures of actual normal everyday people instead of glossy photo models and celebrity lookalikes?

Are these models even able to generate pictures of actual normal everyday people instead of glossy photo models and celebrity lookalikes?

Yes.

For example, take a look at this LoRA which is one of my favorites: https://civitai.com/models/259627/bad-quality-lora-or-sdxl

This, along with a proper model and when prompted properly, will give you photos of people who actually look like real people.

Could it be because conventionally beautiful people are photographed more often and so there’s just more training data?

Also because models in photographs are symetrical, emotionless, softly-lit, and have perfect skin imperfections. Things like age lines, wrinkles, scars, emotional expression, deep shadows and asymetry require actual understanding of human anatomy to draw convincingly.

Yet, the Gen ai image producers have no understanding of anything and draw human anatomy convincingly very often. Yin other words, you're wrong. AI does not need anatomy knowledge, that's not how any of this works. They just need enough training data.

AI does not need anatomy knowledge, that's not how any of this works.

Surely someone has done a paired kinematics model to filter results by this point?

Not my field, but I figured 11 fingered people were just because it was computationally cheaper to have the ape on the other side of the keyboard hit refresh until happy.

There are "normal diffusion" models that create average people with flaws in the style of a low grade consumer camera. They're kind of unsettling because they don't have the same uncanny valley as the typical supermodel photographed with a $10k camera look, but there is still weirdness around the fringes.

Try making a picture of people working in an office with any diffusion model.

It looks like the stock photo cover for that mandatory course you hated.

Even adding keywords like “everyday” doesn’t help. And a fear it’s going to be worse in a few years when this stuff constitutes the majority of the input.

If you prompt the model in the way that pushes it towards low quality amateur photography, the results tend to be more realistic and not glossy:

"90s, single use camera, documentary, of anoffice worker in an open plan office, realistic, amateur photo, blurry"

Results: https://imgur.com/a/GJLqYft

Results: https://imgur.com/a/GJLqYft

Corporate accounts payable, Nina speaking. Just a moooment. https://m.youtube.com/watch?v=4s5yHUpumkY

Is he squinting or am I?

only a prompter would think that looks realistic

Oh a "prompter" is it? No, I'm not a prompter, but if you want to interface with a GenAI model, a prompt is kind of the way to do it. No need to be a salty about it.

Yes, however, for the purposes of a demo article like this, it's significantly easier to use famous people who are essentially baked into the model. You encapsulate all of their specific details in a keyword that's easily reusable across a variety of prompts.

By contrast, a "normal" random person is very easy to generate, but very difficult to keep consistent across scenes.

I know. As well these model outputs are so messed up, it's too much. Penis fingers, vacant cross-eyed stares. This has got to be fine. "Trained on explicit images". It kills me every time. TFA is so emphatic, I think the author is hallucinating as badly as the models are.

Also clothing zippers and seams that end nonsensically.

Not really, I'm in my sixties, and it's surprisingly difficult to get round the biases these models have of young, perfect people, if you want to get images of older people.

Try the single word prompt 'woman' and see what you get...

biases these models have

Have larger diffusion models gotten to synthetic training dogfooding yet?

The irony is that once we get there, we can address biases in historical data. I.e. having a training set that matches reality vs images that were captured and available ~2020.

The larger base models do an excellent job of aging, try out asking for ages increased by 5 year increments, and you’ll see clear progression (some of it caricatured of course), e.g. “55 year old woman” vs “woman”.

so put "old woman" then.

Yeah, I mean it's not great that the models are biased around a certain subset of "woman" (usually young, pretty, white etc.) but you can just describe what you want to see and push the model to give it you. Yes, sometimes it's a bit of a fight, but it's doable.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.