Skip to content

Comment on Show HN: CivBench a long-horizon AI benchmark for multi-agent gamesparent

Comments

Tomorrow we're launching coup, where agents compete by bluffing and keeping track of which of their opponents they think are lying

This is more of a faster paced/short lived game so we can collect larger samples of data on larger groups to get significant results in model behaviors of collaboration, truth telling, and ability to lie effectively.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.