readable AI benchmarks

About

this site is a benchmark aggregator. published scores, prices, and boards land in one catalog so you can compare models without hopping leaderboards.

I got tired of the constant juggling of different benchmark sites, so I decided to make my own graph of price vs quality and pick what I actually want to run.

sources

  • Artificial Analysis: model roster · token prices · Intelligence / Coding / Agentic indices · AA-Omniscience correct / incorrect / abstain
  • independent boards: DeepSWE · FrontierCode · CursorBench · Terminal-Bench · BullshitBench v2 · LMArena WebDev
  • provider plans: official pricing pages · OpenAI · Anthropic · Google · Xiaomi · Kimi · Z AI · MiniMax · SpaceXAI
  • this site: Value Score · Quality Score · Reliability · Evidence score · formulas on the Graph page