Metrics & Ranking Guide

What each score means, how rankings differ, how the two board families compare, and where data comes from. All scores are anchored (against fixed anchors, not daily pool ranks), not absolute.

1. What the scores mean

  • Overall:General impression, good for a first glance.
  • Consensus:Combines 5 benchmark families into one score for general selection; n/5 shows coverage. Each family is first mapped to a fixed anchored score (0–100) and then weighted: new models no longer rescale incumbents, and the leader sits near 90, not 100 (at-anchor per family: AA ≈92, Arena ≈84, Coding ≈92, Agent ≈93, Reasoning ≈90).
  • Coding / Agent / Reasoning:Single-capability scores: coding, automation workflows, deep reasoning.
  • Value:Cheap + long context + smart enough, for tight budgets.
  • Source boards:Raw source-board data:
    • SWE-bench Verified / Pro:Fixing real GitHub bugs, the coding benchmark.
    • Terminal-Bench:Doing tasks in a terminal environment.
    • GPQA / HLE:Graduate-level science QA and hard general exams.
    • Arena ELO:Human-vote arena rankings, overall reputation.
    • OSWorld / BrowseComp:OS operation and web-research agent skills.
  • Speed:Tokens per second, matters for interactive use.

Ranks like (5/1046) after a score mean its position among all models on that metric.

2. How to pick a ranking

  • Primary (Overall / Consensus / Coding / Agent / Reasoning / Value):Pick one for your use case.
  • Benchmark boards / Performance:For single source boards, one at a time.
  • Filters (vendor / open-source / price / capability):Stackable with any ranking; filters rows only. Capability = having that score (same dimensions as tabs); only 16–21% of models have agent/coding scores.

Models missing data sink to the bottom as —, meaning no score yet, not a bad model.

3. Boards: Internal calculations vs OpenRouter source boards

Two families co-exist, physically separated and never intermixed:

  • Internal (All Models sorting):Consensus (AA 30% + Arena 20% + Coding 20% + Agent 15% + Reasoning 15%, anchored per-family scores weighted, ≥2 evidences, renormalized) and 4 single-dimension boards: Coding / Agentic / Reasoning / Value (Value = anchored Pareto of price low + context large + intelligence). All derived, not copied.
  • OpenRouter source boards (/rankings/):Weekly usage (OpenRouter weekly token calls Top10, popularity not capability) and OpenRouter-defined boards: Smartest Open / Best Value / Fastest / Smartest Coding. Ranks are the OpenRouter page order as-is, no re-weighting; each board is prefixed OpenRouter · and links to the original.

Naming guard: OpenRouter · Best Value ≠ internal Value (former is OpenRouter rank, latter is internal Pareto); OpenRouter · Fastest ≠ Speed (former is board rank, latter is tok/s); OpenRouter · Smartest Coding ≠ Coding score (former is original board, latter is coding_index + SWE-bench). The OpenRouter · prefix disambiguates.

Local Coding / Reasoning Top10 also appears at the top of /rankings/ (same source as the All Models tabs, per-group max, computed internally, prefixed Local · ): the prefix separates derived local boards from AA · / OpenRouter · source boards, which reproduce original ranks as-is.

Boards shown are those with data; missing boards auto-hide without breaking the page. Weekly usage lives here, not duplicated in All Models.

See: Rankings → · All Models →

4. Where data comes from

  • Artificial Analysis:Intelligence / coding / agent indexes.
  • LM Arena:Human-vote arena ELO.
  • BenchLM:Task-specific benchmark scores.
  • OpenRouter:Pricing, context windows, weekly usage and OpenRouter board originals.
  • Hugging Face:Open-weights info.
  • Official release pages and blogs.

Sources are documented here instead of per-row labels.

Pricing: official prices first; otherwise other reliable channels (hover shows source); estimates marked ≈. Provider prices may differ; official prices prevail.

Updated daily. Current spec: Index v2. Consensus / Overall / Coding / Agent / Reasoning / Value are anchored scores (absolute, comparable across days); bar widths are relative percentiles (position within the pool), except the Overall bar whose width is its absolute score.

5. Want more boards?

We currently offer source boards (OpenRouter / AA) plus local computed boards (Coding / Reasoning Top10) — if you’d like to see other board types, we welcome your feedback.