Comment on Dola Decoding by Contrasting Layers Improves Factuality in Large Language ModelsComments−totetsu2yexploiting the fact that factual knowledge in an LLMs has generally been shown to be localized to particular transformer layersThis is surprising−dinobones2yWhy is this surprising?It makes sense that “facts” exist in earlier layers, then these become more abstract as you become deeper.This reminds me of residual connections from CNNs and vision.−snthpy2yInteresting. I hadn't considered that before but makes sense.
Comments
This is surprising
Why is this surprising?
It makes sense that “facts” exist in earlier layers, then these become more abstract as you become deeper.
This reminds me of residual connections from CNNs and vision.
Interesting. I hadn't considered that before but makes sense.