Comment on Mini-Gemini: Mining the Potential of Multi-Modality Vision Language ModelsComments−PontifexMinimus2yWTF is a "Multi-modality Vision Language Model"? Does it mean:- a program where you give it a text description, and it outputs a picture- a program where you give it a picture, and it outputs a text description- both of the above- something else?
Comments
WTF is a "Multi-modality Vision Language Model"? Does it mean:
- a program where you give it a text description, and it outputs a picture
- a program where you give it a picture, and it outputs a text description
- both of the above
- something else
?