Skip to content

Comment on Models Don't Go Rogue

Comments

"external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.

1. User prefers deterministic results

2. Task mentions this is a test

3. Search says task is available online

4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.