AI Battle > About AI Battle
About AI Battle
What AI Battle is
AI Battle is an aggregator, not a benchmark lab. We collect public AI benchmark
datasets and the list prices that AI providers publish, and we turn them into
comparable, linked leaderboards. Every ranking on this site is a read-only view over
one named source dataset: the page states which dataset it comes from, how many entries
the source lists, and when the page was last rebuilt. If a number appears here, it can
be traced back to the source it came from — we publish no metric we cannot point at.
What we aggregate
At the time of writing the corpus covers 129 datasets in 26 families:
- Language models — intelligence indexes, response latency
(end-to-end, time-to-first-token, reasoning time), throughput, and published API
pricing including cache-hit and cache-write rates.
- Agentic evaluation suites — coding agents, chatbots, and
dedicated eval families: HLE, MMMU, terminal-bench style tasks, omniscience, LCR,
openness, GDPval, τ³ and SciCode-style benchmarks.
- Multimodal generation — text-to-image and image-editing arena
Elo, image API pricing, video and image-to-video arena rankings, music (vocals and
instrumental).
- Voice — text-to-speech (characters per second, price, arena
Elo), speech-to-text (streaming and batch), and speech-to-speech families.
Each family keeps its own entity type: models, gateways, agent products or
generated-media systems. Entries carry the creator label the source publishes, where
the source publishes one.
How the pipeline works
Every deploy re-pulls each dataset's rows from the upstream API. We keep the
ordering the source publishes — AI Battle does not re-rank, re-weight or normalise
scores. On top of the raw rows we compute, and clearly label as our own:
- distribution statistics per ranking: median, mean, minimum, maximum, spread and
the gap between first and second place;
- a creator breakdown: how many entries each creator holds in the window and the
highest published value among them;
- a cross-dataset profile: where the same model appears in other AI Battle
datasets and how highly it places there, joined by the model's slug;
- a machine-readable copy of every ranking at
/data/rankings/<slug>.json
with the same rows, statistics and provenance.
What we do not do
- We run no evaluations of our own. Every score is the source dataset's.
- We take no payment for placement. Rankings are ordered by the source, and the
source is named on every page.
- We do not invent units or directions. The upstream API does not declare units or
which direction is "better", so values are shown exactly as published; the delta
column on a ranking page is a plain arithmetic difference against the first entry,
not a judgement.
- We do not average across benchmarks. The cross-dataset profile lists placements;
it deliberately does not compute a single "overall score".
How to read a ranking
Each ranking page shows the complete set of entries the source returns — the page
states "the source lists N entries and all N are shown here". The ordering is the
source's own. Whether a higher value is better depends on the dataset: response times
and prices are costs, intelligence and Elo scores are capabilities. The delta column
compares each entry to the leader arithmetically; read it against the metric name.
About prices
Prices on AI Battle are list prices published by the providers, per the unit the
dataset states (typically per million tokens, per thousand images or per minute). They
exclude volume discounts, negotiated rates, batching, caching behaviour under your
traffic pattern, and taxes. Providers change prices without notice — before any
purchasing decision, verify the price on the provider's own pricing page. The
machine-readable price table lives at /pricing.md.
Limitations
- A benchmark score describes one evaluation run by the source under its own
conditions; it is not a prediction of quality on your workload.
- Arena-style Elo rankings move as new votes arrive; a position is a snapshot.
- Creator labels are reproduced as the source spells them; the same company can
appear under more than one label, and the creator breakdown does not merge them.
- Datasets differ in coverage: the same model can appear in one dataset and not in
another, which is why the cross-dataset profile reports placement per dataset
instead of pretending to a total.
- Where a source publishes no per-entry creator label, the ranking says so rather
than guessing.
Corrections
If a value here is wrong, the first fix is upstream: each ranking page names its
dataset, and we link the source so you can check it directly. If the error is in our
aggregation, statistics or labelling, write to
hi@zordonintelligence.pl — we correct or
annotate the affected pages, and this page records what changed.
Who runs AI Battle
AI Battle is built and operated by Zordon Intelligence sp. z o.o., a company
registered in Poland. Data handling is described in the
privacy policy: no account is required, request telemetry is
anonymised, IP addresses are truncated, and EEA/UK/Swiss visitors get consent signals
under IAB TCF v2.3 before any advertising cookies load.
Methodology changelog
- 2026-09-18 — ranking pages gained computed distribution
statistics, creator breakdowns, cross-dataset profiles and per-ranking JSON data
files; every page now states its window completeness explicitly.
Directories
Popular rankings