Comment on Show HN: Text-to-video model from scratch (2 brothers, 2 years, 2B params)parentComments−schopra909OP7moThat all being said, you can just delete the T5 from memory after encoding the text so save on memory.The 2B parameters will take up 4 Gb of memory but activations will be a lot more given size of context windows for video.A 720p 5 second video is roughly 100K tokens of context
Comments
That all being said, you can just delete the T5 from memory after encoding the text so save on memory.
The 2B parameters will take up 4 Gb of memory but activations will be a lot more given size of context windows for video.
A 720p 5 second video is roughly 100K tokens of context