Skip to content

Comment on LoRA vs. Full Fine-Tuning: An Illusion of Equivalenceparent

Comments

How does it compare to partially fine-tuning the model by freezing most of the network beside the last few layers?

Idk but if I was guessing, I would guess that that process would be likely to create intruder dimensions in those layers… but hard to say how impactful that would be. Intuitively I would think it would tend to channel a lot of irrelevant outputs towards the semantic space of the new training data, but idk I how well that intuition would hold up to reality.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.