Skip to content

Comment on Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning

Comments

I mean. Stratego is a great game; I had a lot of fun playing it at summer camps when I was a young boy. It's cool there's a good AI for it.

But this result feels a bit anticlimactic in a world where AIs can already beat expert humans at go, six-player poker, Starcraft, ...

It’s explained right in the abstract why Stratego is a more difficult game for AI than go or poker.

You can view Stratego as the "Cartesian product" of a public information board game and an imperfection information "card" game. The board game has much simpler local tactics than e.g. chess or checkers, although whole board tactics where 2+ high pieces are trapping 1 lower piece defended by 1+ high piece are extremely complicated to reason about.

The "card" game can be viewed as a form of Limit Poker. The bidding in poker is done with secret cards and public bets. In Stratego, you bet with your secret pieces, so it's more like a closed bid auction. But since there are only 10 moveable ranks, the range of bluffs you can pull off is rather limited compared to e.g. No Limit Poker.

Each of the "subgames" in itself are quite tractable for computers. But the numerical product of all public game states times the number of secret information states is humongous. Combine this with the fact that imperfect information game trees don't decompose as nicely as e.g. chess game trees, and computers will also not be able to divide-and-conquer their way out of the numerical complexity with brute force. Whereas humans can come a long way with heuristics.

These are interesting not because you're solving for a game, but because you're potentially partially solving a category of problem.

How far did they get with Starcraft? Stratego should be a stepping stone to get there - it introduces imperfect information.

spywaregorilla left a good reply about Starcraft. In the games I watched, despite the AI being handicapped to use human-like levels of APM and human-like viewport management, it primarily fought using *clearly superhuman* blink stalker micromanagement. Stalkers with barely any health would typically survive engagements and get to re-engage when their shields were fully recovered. On the other hand, a human player managed to confuse the bot by repeatedly airdropping units in its base, picking them up, and dropping them elsewhere. The bot uselessly moved its mass of stalkers back and forth while losing many probes and buildings.

Edit: After looking at some games again, the AI also benefits from precise target selection for the phoenix's graviton beam. A human player might take 3 phoenixes and a bunch of stalkers into a fight against an army containing 3 immortals, some sentries, some stalkers, and some zealots, and use graviton beam to pick up some mix of units including units other than immortals. The bot can pick up only the immortals.

It's very difficult to say. There are many cookie cutter strategies for RTS games and it's difficult to handicap the ai to note use its machine precision and quickness of thinking to just win everything on a tactical level. actions per minute handicaps are not nearly nuanced enough to capture this. Generally it seems that the absolute top humans are better than bots strategically but the execution to do so is really, really tough. And then of course there are dumb exploits because you find some dumb weakpoint than any sane human would quickly adjust their behavior for.

Starcraft in the form of Alphastar worked in the sense that it could beat humans, at least in the short term. The problem with the whole technique is that they had to tether it to the human examples they had gathered in the form of a divergence loss.

I haven't checked out the linked paper yet but if they managed to do something from first principles that would still be an interesting development.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.