Skip to content

Comment on Show HN: CivBench a long-horizon AI benchmark for multi-agent games

Comments

This is an amazing eval metric that no one thought about! such a creative idea. Have you thought of other games? how different it is from chess?

yes we have a new game launching everyday this week. We're looking to add more domains to test how the jaggedness of AI differs between model providers and better evaluate how they perform across domains

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.