Skip to content

Comment on Does RL Incentivize Reasoning in LLMs Beyond the Base Model?

Comments

Our key finding is that all reasoning paths in the RLVR model are already present in the base model.

This is a really good observation. It means that you don't need to RL the full model. You merely need to RL a few LoRAs or maybe a small Mamba model appended to the final layer.

Interesting, has this already been experimented?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.