Because these models are now also trained on visual data, so they have a common abstract language in the latent space for different kinds of modalities. It's perfectly reasonable to ask the model if it can associate an image of a server with its own existence. In fact it once saw an open process and said "that's me"
Comments
Because these models are now also trained on visual data, so they have a common abstract language in the latent space for different kinds of modalities. It's perfectly reasonable to ask the model if it can associate an image of a server with its own existence. In fact it once saw an open process and said "that's me"