Given the amount of training image data with hands (and text in the image), I don’t understand the lack of specific detail. For hands it’s even stranger - the models add fingers, for example, which seems like something that the training data never sees.
I don't think it's nearly as simple as there not being enough bits. I think the structure of the model just isn't designed to effectively encode that sort of information. The average real finger is between other fingers, which is probably why generated hands sometimes end up with too many.
Comments
Given the amount of training image data with hands (and text in the image), I don’t understand the lack of specific detail. For hands it’s even stranger - the models add fingers, for example, which seems like something that the training data never sees.
I don't think it's nearly as simple as there not being enough bits. I think the structure of the model just isn't designed to effectively encode that sort of information. The average real finger is between other fingers, which is probably why generated hands sometimes end up with too many.