lmsys chatbot arena

**LMSYS Chatbot Arena** is the most prominent **open platform** for evaluating and comparing large language models through **live human voting**. Users submit prompts that are answered by two anonymous models side by side, then vote on which response is better — producing a continuously updated **Elo-style leaderboard**. **How It Works** - **Blind Evaluation**: Users enter a prompt, and the system routes it to **two randomly selected models**. Responses appear side by side without revealing which model produced which. - **Human Voting**: Users vote for Response A, Response B, or Tie. This produces a **pairwise preference** judgment. - **Elo Rating**: Votes are aggregated using a **Bradley-Terry model** to compute Elo-style ratings, similar to chess rankings. Models that consistently win against strong opponents earn high ratings. - **Leaderboard**: Publicly accessible at **chat.lmsys.org**, updated with thousands of new votes daily. **Why It Matters** - **Real User Preferences**: Unlike automated benchmarks, the Arena captures what actual users prefer in open-ended conversation — a much more **holistic** signal. - **Diverse Prompts**: Users submit whatever they want — creative writing, coding, reasoning, roleplay, factual questions — covering the full range of LLM use cases. - **Model Diversity**: The Arena hosts dozens of models from different providers, enabling **direct comparison** across the industry. - **Statistical Rigor**: With millions of votes, the rankings are highly statistically significant, with tight confidence intervals. **Key Findings** - Arena rankings often **disagree** with automated benchmarks, revealing that benchmark performance doesn't always translate to user preference. - **Frontier models** (GPT-4, Claude, Gemini) consistently top the leaderboard, but the gap with open-source models has been narrowing. **Developed By** LMSYS (Large Model Systems Organization), a research group at **UC Berkeley** led by researchers including Ion Stoica and the Vicuna team. The Arena has become the de facto standard for **LLM rankings** in the AI community.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account