Skip to content

Comment on DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

Comments

The style matching is interesting, but there's no song structure. There's no identifiable chorus in any of the demo songs.

I find this very surprising, because it's one of things I'd have expected a diffusion model to have a chance of achieving.

I suppose it might be because it's latent diffusion.

That can probably be a style in itself (if we kept exploring in these directions)

No, it’s not a style, it’s by definition an incomplete song.

"Electronic music aren't real songs there's no real instruments involved". Let's be a bit creative with these tools. Sure the pure output isn't always plesant or listenable but there's probably an interesting genre to carve out here

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.