Comment on Does RL Incentivize Reasoning in LLMs Beyond the Base Model?parentComments−zaptrem1yI think when they were figuring out RLHF they avoided this by interleaving RLHF and normal cross entropy on training set gradients.
Comments
I think when they were figuring out RLHF they avoided this by interleaving RLHF and normal cross entropy on training set gradients.