Skip to content

Ask HN: Model sycophancy benchmarks?

5 pointsgradus_ad2 comments
On HN

Model sycophancy is going to become a major issue for individuals and society... Do any benchmarks exist for this? Really needs to become a standard aspect of model quality measures

Comments

Closest I know of are SycEval and the sycophancy evals in Anthropic's 2023 paper, both built on a user pushing back at a correct answer.

I think we should design a "Bullshit Benchmark" that tests for sycophancy and fluff like Opus 5 "two X, but only Y matters" color commentary

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.