Skip to content

Comment on How has DeepSeek improved the Transformer architecture?parent

Comments

The models dwell and draw almost entirely self-referentially from their own reasoning through the entire exchange.

Sounds like a deficiency in theory of the mind.

Maybe explains some of the outputs I've seen from deepseek where it conjectures about the reasons why you said whatever you said. Perhaps this is where we're at for mitigations for what you've noticed.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.