Comment on Does RL Incentivize Reasoning in LLMs Beyond the Base Model?parentComments−Certhas1yIf this was just the effect you mention you would not expect the base model to surpass the RL model though. Plus their k are much smaller than that.I think it's a very interesting and meaningful study.
Comments
If this was just the effect you mention you would not expect the base model to surpass the RL model though. Plus their k are much smaller than that.
I think it's a very interesting and meaningful study.