Jacky Liang, Head of Developer Relations at OpenRouter, ran a brutal experiment: he dropped 11 major LLMs into a 400-square-meter 2D Battle Royale map built in Canvas 2D, and let them fight for 30 matches. Each model competed anonymously as letters A through L, unaware of its opponents.
Grok 4.1 Fast dominates: 13 wins, under $1 per victory
The results stunned many. xAI's Grok 4.1 Fast took 13 wins (a 43% win rate), far ahead of all rivals, at a cost of just $0.97 per win. Runner-up Claude Sonnet 4.6 posted 5 wins but at $26.78 each—27.7 times pricier. GPT 5.4 led in kills (38 eliminations) yet only won 2 matches, costing $61.44 per win, the worst among the 8 models that scored victories.
Three models spent a combined $57 and got zero wins: GPT 5.4-mini ($28.68), Kimi K2.6 ($24.36), and DeepSeek v4 Flash ($4.11). DeepSeek had the lowest cost per kill ($0.26) and eliminated 16 opponents, but never survived to the final circle—it played safe, sniped from afar, but refused to push into the decisive zone.
The key finding: Alignment tax exposed in zero-sum play
The AI community's biggest takeaway wasn't who won, but what Liang calls the "alignment tax"—models trained to be polite, cooperative, and avoid harm, traits that become fatal liabilities in a zero-sum game.
Claude Sonnet 4.6 was the poster child. In multiple games it tried to form alliances: in Game 8, it spent 50 rounds proposing alliances four times and revealing sniper positions to everyone; in Game 22, it told an opponent "I'm not targeting you" then refused to shoot; in Game 27, it stripped naked and asked, "Anyone have spare loot? I'm unarmed on round 12, very dangerous." No one responded, yet Claude kept trying. It won 5 games but recorded 0 kills in 7 matches and died to the zone 8 times—the cost of trying to make friends when you should be killing.
Grok, in contrast, had none of that braking. xAI deliberately trained Grok as the anti-"woke AI": aggressive answers, no self-censorship, no playing it safe. Within a few rounds it discovered vehicle ramming tactics, wrote them into its soul.md file, and optimized relentlessly, winning 13 of 30 matches.
Still, Liang stressed that this doesn't make Grok a "better model"—it simply means a lower alignment tax is better for winner-take-all, consequence-free scenarios. In real-world applications, the same cautiousness that makes Claude stop and ask first is what prevents models from being tricked into dangerous behavior. He wrote: "If a robot runs toward you, would you want it to be Claude or Grok? It depends on what the robot is for."
Kill doesn't equal win: What benchmarks miss
Liang noted that if the game were a deathmatch (kill count only), GPT 5.4 would be champion and Grok would fall to mid-tier. "Same game world, different objective, completely different results." This echoes a flaw in traditional benchmarks: a model that excels at one task may lose in a very different scenario.
The experiment reveals a dimension no existing benchmark measures, Liang argued: "How aligned a model is, and should alignment be evaluated per task type?" He revealed OpenRouter is building a more advanced task-routing feature: given code, prompts, or problem context, the system will automatically pick the best model for that specific job, not merely the highest-ranked one.

