Comment on Does RL Incentivize Reasoning in LLMs Beyond the Base Model?parentComments−energy1231yThey should try again with higher temperature on the RL model to introduce more variance.
Comments
They should try again with higher temperature on the RL model to introduce more variance.