Comment on Does RL Incentivize Reasoning in LLMs Beyond the Base Model?Comments−whatshisface1yIf you don't know the answer to a problem, you're not going to be able to repeat sampling until it is correct. Random strings will saturate all benchmarks at k=infinity if tested this way.
Comments
If you don't know the answer to a problem, you're not going to be able to repeat sampling until it is correct. Random strings will saturate all benchmarks at k=infinity if tested this way.