Comment on Show HN: CivBench a long-horizon AI benchmark for multi-agent gamesComments−zimbo636moThis is an amazing eval metric that no one thought about! such a creative idea. Have you thought of other games? how different it is from chess?−mbh159OP6moyes we have a new game launching everyday this week. We're looking to add more domains to test how the jaggedness of AI differs between model providers and better evaluate how they perform across domains
Comments
This is an amazing eval metric that no one thought about! such a creative idea. Have you thought of other games? how different it is from chess?
yes we have a new game launching everyday this week. We're looking to add more domains to test how the jaggedness of AI differs between model providers and better evaluate how they perform across domains