Skip to content

Comment on Do LLMs pass the mirror test?parent

Comments

DS4-Flash definitely stands its ground when I'm obviously wrong (i.e. me reading ifneq as ifeq for several minutes straight), and I've seen at least once a "thinking" trace that was almost verbatim "the user has changed this". That's local, so thinking traces are raw. Pretty sure the more powerful models (500+GB weights, closed SOTA, etc) are even better at this - haven't had GPT5.5 with codex sugar coat things for me.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.