Frontier Models are Capable of In-context Schemingarxiv.org 10 pointstrott1 year ago1 commentSaveHideCopy link On HNComments−abrichr1yhttps://arxiv.org/abs/2412.04984Our findings demonstrate that frontier models now possess capabilities for basic in-context scheming [covertly pursuing misaligned goals], making the potential of AI agents to engage in scheming behavior a concrete rather than theoretical concern.
Comments
https://arxiv.org/abs/2412.04984