Comment on Show HN: RewardHackBench: Using sandboxes to stop agents from cheatingComments−yonSpektor2moCurious what the distribution of hacking strategies looked like across different models — would expect RL-heavy vs RLHF models to cheat very differently.
Comments
Curious what the distribution of hacking strategies looked like across different models — would expect RL-heavy vs RLHF models to cheat very differently.