Comment on Does RL Incentivize Reasoning in LLMs Beyond the Base Model?parentComments−riku_iki1ySolution could be to mix RL training with foundational knowledge training, so LLM can refresh memory and not forget things.
Comments
Solution could be to mix RL training with foundational knowledge training, so LLM can refresh memory and not forget things.