Skip to content

Comment on Continuous Diffusion Language Models (CDLM's)parent

Comments

I suppose what I was trying to say is that in the research community, people were still a lot more willing to entertain alternative modelling paradigms for language at that point, there was much less of a monoculture than there is today. It was really only after ChatGPT that non-autoregressive language modelling came to be seen as a fringe pursuit. Of course that is a highly subjective assessment, and my perspective is inevitably coloured by my own research interests at the time, and those of the people around me.

For what it's worth, I don't believe post-training diffusion models (of language or otherwise) is uniquely difficult, just a lot less thoroughly explored so far.

That’s fair, but my impression is that AI is a very industry-driven field. When commercialization fails, industry funding dries up and drags gov funding down too, resulting in an AI winter.

And industry really cares about solving practical problems. So as soon as autoregression “solved” the text generation problem (albeit imperfectly) the field was ready to move on. If text diffusion is not faster, cheaper, or higher quality, there is no ROI.

In fact text diffusion is much faster, so I have used it in the past. And I think it has the potential to be better as well, given a constant time budget. So I think it can be a really promising method moving forward.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.