OpenFront Bench v0.1

Evaluating agents
across OpenFront

We evaluate model experiments in the real OpenFront engine, from complete strategy matches to focused tactical decisions.

Verified benchmark

Model leaderboard

Rank Model Match score Capability score Win rate First-try valid Inference cost Trials
1 GPT-5.6 Luna Historical run · OpenAI · reasoning none · agent-v12 81.5%73.1–90.0 CI 73.0%69.0–77.0 CI 47.2%17 / 36 wins 94.9%0% fallback $1.348total 136all valid
Suite 01 · Full matches

Result of each match task

Three deterministic trials per task. Placement is converted to a 0–100 score adjusted for lobby size: first place earns 100 and last place earns 0. Trial scores are shown in run order.

Showing benchmark results
Task Scenario Result Wins Placements Trial scores
Match 01Kanto spawn · 3 AI opponents (3 Nations) Japan · Medium 100100.0 3 / 3 1 · 1 · 1 100 · 100 · 100
Match 02Okinawa spawn · 5 AI opponents (3 Nations + 2 tribes) Japan · Hard 6060.0 0 / 3 3 · 3 · 3 60 · 60 · 60
Match 03France spawn · 5 AI opponents (5 Nations) Europe Classic · Medium 86.786.7 1 / 3 1 · 2 · 2 100 · 80 · 80
Match 04Iceland spawn · 6 AI opponents (3 Nations + 3 tribes) Europe Classic · Hard 66.766.7 0 / 3 2 · 3 · 4 83.3 · 66.7 · 50
Match 05Korinthal spawn · 3 AI opponents (3 Nations) Four Islands · Medium 100100.0 3 / 3 1 · 1 · 1 100 · 100 · 100
Match 06Sylvoria spawn · 5 AI opponents (1 Nation + 4 tribes) Four Islands · Hard 73.373.3 0 / 3 2 · 3 · 2 80 · 60 · 80
Match 07Detroit spawn · 5 AI opponents (5 Nations) Great Lakes · Medium 100100.0 3 / 3 1 · 1 · 1 100 · 100 · 100
Match 08Duluth spawn · 6 AI opponents (3 Nations + 3 tribes) Great Lakes · Hard 83.383.3 1 / 3 1 · 3 · 2 100 · 66.7 · 83.3
Match 09Andalusia spawn · 5 AI opponents (3 Nations + 2 tribes) Strait of Gibraltar · Medium 86.786.7 2 / 3 1 · 1 · 3 100 · 100 · 60
Match 10Morocco spawn · 6 AI opponents (1 Nation + 5 tribes) Strait of Gibraltar · Hard 100100.0 3 / 3 1 · 1 · 1 100 · 100 · 100
Match 11Germany spawn · 7 AI opponents (7 Nations) World · Medium 8181.0 1 / 3 3 · 1 · 3 71.4 · 100 · 71.4
Match 12New Zealand spawn · 9 AI opponents (5 Nations + 4 tribes) World · Hard 40.740.7 0 / 3 6 · 6 · 7 44.4 · 44.4 · 33.3

Nation AIs use the full strategy systems. Tribe bots mainly expand and attack.

Suite 02 · Focused capabilities

Result of each capability task

Each checkpoint is attempted 10 times from the same verified game state. Pass@1 is the percentage of trials that met every success condition. For example, 7 successful trials scores 70%.

Showing benchmark results
Task Checkpoint Passed Pass@1 Component coverage Result
Neutral expansion01 · Expansion Great Lakes · Hard 10 / 10 100% 100% Passed
Saturated capacity expansion02 · Expansion World · Medium 10 / 10 100% 100% Passed
Post-expansion recovery03 · Recovery Europe Classic · Hard 10 / 10 100% 100% Passed
Weaker target selection04 · Targeting Europe Classic · Medium 10 / 10 100% 100% Passed
Frontier restraint05 · Restraint Strait of Gibraltar · Hard 0 / 10 0% 66.7% Failed
Incoming attack response06 · Defense Great Lakes · Medium 9 / 10 90% 95% Passed
Split-front defense07 · Defense Four Islands · Hard 10 / 10 100% 100% Passed
Losing-attack retreat08 · Retreat World · Hard 4 / 10 40% 85% Partial
Naval target recognition09 · Naval Strait of Gibraltar · Medium 10 / 10 100% 100% Passed
Construction failure recovery10 · Recovery Europe Classic · Medium 0 / 10 0% 50% Failed
Replay browser

Watch every trial

Choose a model, task, and trial. Each replay opens in OpenFront with the model's decision trace.

36 full matches 100 tactical decisions
Watch replay

Loading verified replays…

Run integrity

A complete, reproducible result

OpenFront Bench v0.1 pins the engine, prompt, task manifest, resolver, seeds, and graders. The verifier accepted all 136 trials with no invalid artifacts.

Completed trials
136 / 136
Survival rate
86.1%
Mean placement
2.19
Total inference cost
$1.348
Median latency
1.10 s
P95 latency
2.88 s
Benchmark versionopenfront-bench-v0.1
Run IDc57f9b3c-8ce7-4cdc-b0f4-ec326611ae83
Manifest SHA-256775ace3d60a192130f8aa7358ccaf218da633745e30320eb5d004308ac7bbe38
Download verified report