OpenGameEval tests agents in Roblox Studio, finding they still struggle with multi-game creation tasks.
By GameSlash
Tested 13 models across 84 primary tasks in Roblox Studio, evaluating both success rates and scene exploration, with test suites and plugins released.
OpenGameEval deploys language agents in Roblox Studio and evaluates both the generated scenes and the outcomes during simulated gameplay. The research was published on arXiv on October 1, 2026, to assess game creation in continuous-state environments, rather than relying solely on code or final outputs.

The team tested 13 models on 84 human-curated main tasks, with 16 attempts per task. The best model succeeded on 51.7% of tasks in a single attempt, and achieved success in all five of five attempts on 39.4%. Six tasks remained unsolved by every tested model.
Inspecting the scene before editing improves the chances of success
The team's report found that agents who inspected related objects before making changes had a higher pass rate. Beyond the final results, this test framework also captures in-process behavior to determine whether agents retrieve all necessary information.
The team released the task set, scene files, report documentation, plugin, and results table under MIT, allowing developers to study or run further tests. The reported scores apply specifically to this task set.
Data and test kits: OpenGameEval on arXiv · Roblox / open-game-eval on GitHub