Comment on Text-to-4D Dynamic Scene GenerationComments−radarsat13ytrained only on Text-Image pairs and unlabeled videosThis is fascinating. It's able to pick up sufficiently on the fundamentals of 3D motion from 2D videos, while only needing static images with descriptions to infer semantics.
Comments
This is fascinating. It's able to pick up sufficiently on the fundamentals of 3D motion from 2D videos, while only needing static images with descriptions to infer semantics.