Back to AI game news
AI game news

AAArena tests AI agents writing bots for 12 games, with No. 1 results in 6 games

By GameSlash

Tsinghua's AAArena research uses 12 games and 1,920 original bot sets to test AI that reads rules, writes bots, and adjusts strategies from replays. The best results ranked first in 6 games.

AIAI AgentResearch

Tsinghua University research team published AAArena, which is a test suite for AI agents on October 8, 2026. It asks AI to read the rules, write programs to play games, and learn from competition results across 12 games. The reported results show that Claude Opus 5.5 paired with Claude Code took a bot to rank 1 in 6 games, compared with the archived competitor programs.

AAArena diagram shows 12 competitive games and Elo comparison graphs of bots from multiple AI systems versus original competitor programs.
Figure 1 from AAArena by Kaisen Yang et al. (2026), arXiv:2610.12341 — CC BY 4.0

Image: Original image in the research article | CC BY 4.0 License

AAArena does not allow the model to press character control buttons at every moment, but instead enables an agent to develop the code for a bot used in actual competitions. Upon defeat, the agent reads match logs and adjusts strategies in the code to retest, all without modifying the weights of the underlying AI model.

The team compiled 1,920 sets of competitor programs from human competitions, spanning multiple genres including maze navigation, territory conquest, and resource management. This figure of 1,920 refers to the number of programs in the comparison database, not the number of unique human players. The comparative experiment involved 7 agent types under a limited budget for competition trials and ranking evaluation.

Although the best system ranked No. 1 in half the games, none of the tested systems reached first place in the other 6 games. The research suggests that complex rulebooks make it harder for bots to develop strategies. These results are not evidence that AI can play games or write bots better than humans in every game.

The results highlight the difference between AI that merely answers questions and AI Agents that adapt their own software based on actual outcomes. Read the work from another perspective at OpenGameEval testing agents in Roblox Studio which tests game creation rather than competitive bot programming.

Source: AAArena Research Article (October 8, 2026) | Project Website | Code on GitHub

View source