Evaluating agents
across OpenFront
We evaluate model experiments in the real OpenFront engine, from complete strategy matches to focused tactical decisions.
Model leaderboard
| Rank | Model | Match score | Capability score | Win rate | First-try valid | Inference cost | Trials |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna Historical run · OpenAI · reasoning none · agent-v12 | 81.5%73.1–90.0 CI | 73.0%69.0–77.0 CI | 47.2%17 / 36 wins | 94.9%0% fallback | $1.348total | 136all valid |
Result of each match task
Three deterministic trials per task. Placement is converted to a 0–100 score adjusted for lobby size: first place earns 100 and last place earns 0. Trial scores are shown in run order.
| Task | Scenario | Result | Wins | Placements | Trial scores |
|---|---|---|---|---|---|
| Match 01Kanto spawn · 3 AI opponents (3 Nations) | Japan · Medium |
|
3 / 3 | 1 · 1 · 1 | 100 · 100 · 100 |
| Match 02Okinawa spawn · 5 AI opponents (3 Nations + 2 tribes) | Japan · Hard |
|
0 / 3 | 3 · 3 · 3 | 60 · 60 · 60 |
| Match 03France spawn · 5 AI opponents (5 Nations) | Europe Classic · Medium |
|
1 / 3 | 1 · 2 · 2 | 100 · 80 · 80 |
| Match 04Iceland spawn · 6 AI opponents (3 Nations + 3 tribes) | Europe Classic · Hard |
|
0 / 3 | 2 · 3 · 4 | 83.3 · 66.7 · 50 |
| Match 05Korinthal spawn · 3 AI opponents (3 Nations) | Four Islands · Medium |
|
3 / 3 | 1 · 1 · 1 | 100 · 100 · 100 |
| Match 06Sylvoria spawn · 5 AI opponents (1 Nation + 4 tribes) | Four Islands · Hard |
|
0 / 3 | 2 · 3 · 2 | 80 · 60 · 80 |
| Match 07Detroit spawn · 5 AI opponents (5 Nations) | Great Lakes · Medium |
|
3 / 3 | 1 · 1 · 1 | 100 · 100 · 100 |
| Match 08Duluth spawn · 6 AI opponents (3 Nations + 3 tribes) | Great Lakes · Hard |
|
1 / 3 | 1 · 3 · 2 | 100 · 66.7 · 83.3 |
| Match 09Andalusia spawn · 5 AI opponents (3 Nations + 2 tribes) | Strait of Gibraltar · Medium |
|
2 / 3 | 1 · 1 · 3 | 100 · 100 · 60 |
| Match 10Morocco spawn · 6 AI opponents (1 Nation + 5 tribes) | Strait of Gibraltar · Hard |
|
3 / 3 | 1 · 1 · 1 | 100 · 100 · 100 |
| Match 11Germany spawn · 7 AI opponents (7 Nations) | World · Medium |
|
1 / 3 | 3 · 1 · 3 | 71.4 · 100 · 71.4 |
| Match 12New Zealand spawn · 9 AI opponents (5 Nations + 4 tribes) | World · Hard |
|
0 / 3 | 6 · 6 · 7 | 44.4 · 44.4 · 33.3 |
Nation AIs use the full strategy systems. Tribe bots mainly expand and attack.
Result of each capability task
Each checkpoint is attempted 10 times from the same verified game state. Pass@1 is the percentage of trials that met every success condition. For example, 7 successful trials scores 70%.
| Task | Checkpoint | Passed | Pass@1 | Component coverage | Result |
|---|---|---|---|---|---|
| Neutral expansion01 · Expansion | Great Lakes · Hard | 10 / 10 | 100% | 100% | Passed |
| Saturated capacity expansion02 · Expansion | World · Medium | 10 / 10 | 100% | 100% | Passed |
| Post-expansion recovery03 · Recovery | Europe Classic · Hard | 10 / 10 | 100% | 100% | Passed |
| Weaker target selection04 · Targeting | Europe Classic · Medium | 10 / 10 | 100% | 100% | Passed |
| Frontier restraint05 · Restraint | Strait of Gibraltar · Hard | 0 / 10 | 0% | 66.7% | Failed |
| Incoming attack response06 · Defense | Great Lakes · Medium | 9 / 10 | 90% | 95% | Passed |
| Split-front defense07 · Defense | Four Islands · Hard | 10 / 10 | 100% | 100% | Passed |
| Losing-attack retreat08 · Retreat | World · Hard | 4 / 10 | 40% | 85% | Partial |
| Naval target recognition09 · Naval | Strait of Gibraltar · Medium | 10 / 10 | 100% | 100% | Passed |
| Construction failure recovery10 · Recovery | Europe Classic · Medium | 0 / 10 | 0% | 50% | Failed |
Watch every trial
Choose a model, task, and trial. Each replay opens in OpenFront with the model's decision trace.
A complete, reproducible result
OpenFront Bench v0.1 pins the engine, prompt, task manifest, resolver, seeds, and graders. The verifier accepted all 136 trials with no invalid artifacts.
- Completed trials
- 136 / 136
- Survival rate
- 86.1%
- Mean placement
- 2.19
- Total inference cost
- $1.348
- Median latency
- 1.10 s
- P95 latency
- 2.88 s
openfront-bench-v0.1
c57f9b3c-8ce7-4cdc-b0f4-ec326611ae83
775ace3d60a192130f8aa7358ccaf218da633745e30320eb5d004308ac7bbe38