Skip to content

Comment on Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learningparent

Comments

How far did they get with Starcraft? Stratego should be a stepping stone to get there - it introduces imperfect information.

spywaregorilla left a good reply about Starcraft. In the games I watched, despite the AI being handicapped to use human-like levels of APM and human-like viewport management, it primarily fought using *clearly superhuman* blink stalker micromanagement. Stalkers with barely any health would typically survive engagements and get to re-engage when their shields were fully recovered. On the other hand, a human player managed to confuse the bot by repeatedly airdropping units in its base, picking them up, and dropping them elsewhere. The bot uselessly moved its mass of stalkers back and forth while losing many probes and buildings.

Edit: After looking at some games again, the AI also benefits from precise target selection for the phoenix's graviton beam. A human player might take 3 phoenixes and a bunch of stalkers into a fight against an army containing 3 immortals, some sentries, some stalkers, and some zealots, and use graviton beam to pick up some mix of units including units other than immortals. The bot can pick up only the immortals.

It's very difficult to say. There are many cookie cutter strategies for RTS games and it's difficult to handicap the ai to note use its machine precision and quickness of thinking to just win everything on a tactical level. actions per minute handicaps are not nearly nuanced enough to capture this. Generally it seems that the absolute top humans are better than bots strategically but the execution to do so is really, really tough. And then of course there are dumb exploits because you find some dumb weakpoint than any sane human would quickly adjust their behavior for.

Starcraft in the form of Alphastar worked in the sense that it could beat humans, at least in the short term. The problem with the whole technique is that they had to tether it to the human examples they had gathered in the form of a divergence loss.

I haven't checked out the linked paper yet but if they managed to do something from first principles that would still be an interesting development.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.