Skip to content

Comment on The VAE Used for Stable Diffusion Is Flawed

Comments

I'm curious, was this well-known by experts already? How surprising is this?

I enjoyed the write up.

If one ever tried to make edits to the latents prior to decoding them with a VAE in SD1.5 and then in SDXL, it could be seen that that local changes had somewhat unpredictable and global effects on the image in SD1.5, while in SDXL the changes have more predictable impacts to the output image and some of the different latent channels end up corresponding more directly to the resulting image channels.

Definitely a fascinating write-up. I have been curious about these differences for a while, though I had never considered this a "problem" per se.

I have never heard of this problem before, and I have seen a lot of discussion about VAE from researchers.

I've once seen someone on Twitter wondering about to-them-obviously-bug with VAE leading to oddly saturated images in anime space, just my dumb brain keyword search though

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.