Skip to content

Comment on Text-to-4D Dynamic Scene Generation

Comments

trained only on Text-Image pairs and unlabeled videos

This is fascinating. It's able to pick up sufficiently on the fundamentals of 3D motion from 2D videos, while only needing static images with descriptions to infer semantics.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.