Comment on Dola Decoding by Contrasting Layers Improves Factuality in Large Language ModelsparentComments−dinobones2yWhy is this surprising?It makes sense that “facts” exist in earlier layers, then these become more abstract as you become deeper.This reminds me of residual connections from CNNs and vision.−snthpy2yInteresting. I hadn't considered that before but makes sense.
Comments
Why is this surprising?
It makes sense that “facts” exist in earlier layers, then these become more abstract as you become deeper.
This reminds me of residual connections from CNNs and vision.
Interesting. I hadn't considered that before but makes sense.