LMArena
Crowdsourced blind arena where humans vote on anonymous model responses.
why this verdict
Keep it — OpenAI has not replaced it
OpenAI's version overlaps, but does not finish the research job, so this one is still worth keeping open.
- A named launch, not a vibe named, unlinked
- Labs publishing their own benchmark results and independent evaluation suites multiplying. OpenAI, August 5, 2025 — no announcement link recorded yet.
- How much of the job it covers not the job editorial call
- Parts of it. The job still needs the tool to get finished.
- Is there a free way to do it? yes
- 3 of 3 listed replacements have a usable free tier: Artificial Analysis, SWE-bench and Hugging Face.
- What the call is worth nothing to cancel
- No paid entry tier tracked, so there is no subscription to cancel.
- Threatened by
- OpenAI
- Since
- August 5, 2025
- List price
- free
- Per year
- —
The backstory
LMArena, which grew out of Chatbot Arena at Berkeley, ranks models by having people compare two anonymous answers and pick the better one, then computing Elo ratings from millions of votes. It became the scoreboard labs cite in launch posts, which is enormous influence for a research project. It has been criticised for rewarding pleasant formatting over correctness, and the arrival of many task-specific benchmarks dilutes it, but no other public ranking carries the same weight.
Escape hatches
Independent benchmarks for speed, price and quality
artificialanalysis.ai open_in_newTask-specific benchmark for real software engineering
swebench.com open_in_newOpen leaderboards across many evaluation suites
huggingface.co open_in_new3 of 3 replacements have a usable free tier.