lmsys chatbot arena
**LMSYS Chatbot Arena** is the most prominent **open platform** for evaluating and comparing large language models through **live human voting**. Users submit prompts that are answered by two anonymous models side by side, then vote on which response is better — producing a continuously updated **Elo-style leaderboard**.
**How It Works**
- **Blind Evaluation**: Users enter a prompt, and the system routes it to **two randomly selected models**. Responses appear side by side without revealing which model produced which.
- **Human Voting**: Users vote for Response A, Response B, or Tie. This produces a **pairwise preference** judgment.
- **Elo Rating**: Votes are aggregated using a **Bradley-Terry model** to compute Elo-style ratings, similar to chess rankings. Models that consistently win against strong opponents earn high ratings.
- **Leaderboard**: Publicly accessible at **chat.lmsys.org**, updated with thousands of new votes daily.
**Why It Matters**
- **Real User Preferences**: Unlike automated benchmarks, the Arena captures what actual users prefer in open-ended conversation — a much more **holistic** signal.
- **Diverse Prompts**: Users submit whatever they want — creative writing, coding, reasoning, roleplay, factual questions — covering the full range of LLM use cases.
- **Model Diversity**: The Arena hosts dozens of models from different providers, enabling **direct comparison** across the industry.
- **Statistical Rigor**: With millions of votes, the rankings are highly statistically significant, with tight confidence intervals.
**Key Findings**
- Arena rankings often **disagree** with automated benchmarks, revealing that benchmark performance doesn't always translate to user preference.
- **Frontier models** (GPT-4, Claude, Gemini) consistently top the leaderboard, but the gap with open-source models has been narrowing.
**Developed By**
LMSYS (Large Model Systems Organization), a research group at **UC Berkeley** led by researchers including Ion Stoica and the Vicuna team. The Arena has become the de facto standard for **LLM rankings** in the AI community.