Skip to content

Comment on Unlocking a Million Times More Data for AI

Comments

I remember that a couple of years ago people were talking about how multimodal models would have skills bleed-over, so one that's trained on the same amount of text + a ton of video/image data would perform better on text responses. Did this end up holding up? Intuitively I would think that text packs much more meaning into the same amount of data than visuals do (a single 1000x1000px image would be about the same amount of data as a million characters), which would hamstring it.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.