Skip to content

Comment on Dola Decoding by Contrasting Layers Improves Factuality in Large Language Models

Comments

exploiting the fact that factual knowledge in an LLMs has generally been shown to be localized to particular transformer layers

This is surprising

Why is this surprising?

It makes sense that “facts” exist in earlier layers, then these become more abstract as you become deeper.

This reminds me of residual connections from CNNs and vision.

Interesting. I hadn't considered that before but makes sense.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.