They don't need to train on your photos to make your photos do these things. It's a diffusion technique / model that can operate on any photo it hasn't seen before. I've seen many "make these people smile in the photo" apps. There is also multiple papers on turning photos into videos.
They demoed the photos history / storyboard et al during Google i/o a week or two ago. The "showing the progress of your child learning to swim" was pretty impressive because it included photos of swimming certificates along side ordering the actual swimming photos, all without needing to organize the photos manually, so there is a query/embedding aspect to it
Have you seen image blend / target modification results? Midjourney features are just the tip of the iceberg
There is no way Google is training on every photo that it does this to. It is prohibitively expensive and not necessary to get the results you describe. They can just feed in a set of images, with one to be processed and the others as reference, to an already trained model
Comments
They don't need to train on your photos to make your photos do these things. It's a diffusion technique / model that can operate on any photo it hasn't seen before. I've seen many "make these people smile in the photo" apps. There is also multiple papers on turning photos into videos.
They demoed the photos history / storyboard et al during Google i/o a week or two ago. The "showing the progress of your child learning to swim" was pretty impressive because it included photos of swimming certificates along side ordering the actual swimming photos, all without needing to organize the photos manually, so there is a query/embedding aspect to it
I would agree with you, except that it knew exactly what her smile looked like. It animated a photo of her not smiling, into one of her actual smile.
The only way to get that information is from other photos of her smiling.
Have you seen image blend / target modification results? Midjourney features are just the tip of the iceberg
There is no way Google is training on every photo that it does this to. It is prohibitively expensive and not necessary to get the results you describe. They can just feed in a set of images, with one to be processed and the others as reference, to an already trained model