(1) [Hands] are intricately structured, having many subcomponents which have precise spatial interrelationships over a range of scales; there are a lot of ways to make things that are like text/hands except wrong 2) The average person is intimately familiar with said structures, having spent thousands of hours engaging looking at them while performing complex tasks involving a visiospatial feedback loop.
Shouldn't this apply even more strongly to faces versus hands? AI seems to have a significantly easier time with those.
Faces don't have lots of repeating similar subcomponents beyond some things that are just two items in bilateral symmetry (teeth are a big exception, and teeth, when visible, can be a problem.)
And, actually, faces still, especially outside of closeups of just the face, can be a problem, too, which is why a separate face restoration with a GAN or inpainting pass for faces with the same or different diffusion model is common.
faces are simpler since unlike hands most of their major constituents are at fixed relative positions to each other. but the flip side is that people are hyper-biased towards attending to facial details, hence why they were basically the first-handled special case
Faces are probably vastly over-represented in the training data. Normal people and professional photographers alike love photographing faces, at zoom levels that are quite rare to experience in real life.
Comments
Shouldn't this apply even more strongly to faces versus hands? AI seems to have a significantly easier time with those.
Faces don't have lots of repeating similar subcomponents beyond some things that are just two items in bilateral symmetry (teeth are a big exception, and teeth, when visible, can be a problem.)
And, actually, faces still, especially outside of closeups of just the face, can be a problem, too, which is why a separate face restoration with a GAN or inpainting pass for faces with the same or different diffusion model is common.
faces are simpler since unlike hands most of their major constituents are at fixed relative positions to each other. but the flip side is that people are hyper-biased towards attending to facial details, hence why they were basically the first-handled special case
Faces are probably vastly over-represented in the training data. Normal people and professional photographers alike love photographing faces, at zoom levels that are quite rare to experience in real life.