By Zordon Intelligence. Published . Every number below is computed by AI Battle from public datasets refreshed on every deploy; the methodology is on the about page.
The twelve models that appear in the largest number of AI Battle datasets. A wider
footprint does not mean better — it means more independent measurements, which is what
makes a profile at /models/<slug> worth reading.
| Model | Datasets | Best placement |
|---|---|---|
| Muse Spark 1.3 (max) — Meta | 61 | #2 of 11 |
| Grok 4.6 (high) — SpaceXAI | 61 | #2 of 20 |
| Gemini 3.8 Flash (high) — Google | 60 | #1 of 11 |
| Kimi K3 (max) — Kimi | 60 | #1 of 20 |
| DeepSeek V4 Pro 0813 (Reasoning, Max Effort) — DeepSeek | 59 | #3 of 14 |
| GPT-5.6 Luna (max) — OpenAI | 58 | #1 of 11 |
| MiniMax-M3 — MiniMax | 57 | #1 of 20 |
| GLM-5.3 (max) — Z AI | 57 | #3 of 20 |
| DeepSeek V4.1 Flash (Reasoning, Max Effort) — DeepSeek | 56 | #2 of 11 |
| GPT-5.6 Sol (max) — OpenAI | 56 | #3 of 20 |
| GLM-5.3-Flash — Z AI | 54 | #1 of 20 |
| Inkling (xhigh) — Thinking Machines | 52 | #3 of 20 |
These are the models that the source datasets themselves rank highest in three or more independent measurements — capability scores, response times and arena Elo count separately.
| Model | Datasets | Best placement |
|---|---|---|
| Gemini 3.8 Flash (high) — Google | 60 | #1 of 11 |
| Kimi K3 (max) — Kimi | 60 | #1 of 20 |
| GPT-5.6 Luna (max) — OpenAI | 58 | #1 of 11 |
| MiniMax-M3 — MiniMax | 57 | #1 of 20 |
| GLM-5.3-Flash — Z AI | 54 | #1 of 20 |
Open any model's profile for the full placement history; for a head-to-head with another model, use the prerendered pages under /vs or pick any two in the interactive comparison.
Generated automatically from public data. No editorial picks, no paid placements. Source for every metric: the dataset named on its ranking page. Corrections: hi@zordonintelligence.pl.