Skip to content

Comment on The Little Book of Reinforcement Learningparent

Comments

GRPO is policy gradient/PPO with your value function baseline monte carlo estimated using k rollouts. The only new thing is finding out it works well with binary rewards and LLM policies.

It is a huge improvement to PPO because you don’t need a separate critic model which cuts memory costs in half and stabilizes training.

Yes, but monte carlo estimating the critic model is not new.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.