It’s both. Fine-tuning a BERT model changes its weights, which causes the embeddings to change.
For example you might have one model which embeds a text query and another model which embeds an image. You also have a dataset of image + text captions. Training means updating the weights of those models so that the embedding of the image is close to the embedding of its corresponding caption.
Comments
GPT2 and 3 work via autoregression. In chat bots, decoders work via autoregression.
Yes, I was one of them. That’s not called “post-training” it’s called fine-tuning.
Asking out of curiosity, because I have limited experience in that domain.
I thought that fine-tuning was changing the weights in the model, not the embedding? Or did I misunderstand?
It’s both. Fine-tuning a BERT model changes its weights, which causes the embeddings to change.
For example you might have one model which embeds a text query and another model which embeds an image. You also have a dataset of image + text captions. Training means updating the weights of those models so that the embedding of the image is close to the embedding of its corresponding caption.