Comment on OpenAI: support for Reinforcement Fine-tuning available to verified orgsComments−justanotheratomOP1ymy question for anyone who knows:Between SFT, DPO, and RFT, - when to use which? - can we mix and match? e.g, first SFT, then DPO.
Comments
my question for anyone who knows:
Between SFT, DPO, and RFT, - when to use which? - can we mix and match? e.g, first SFT, then DPO.