Home Knowledge Base LLM Leaderboards and Rankings

LLM Leaderboards and Rankings

Major Leaderboards

Chatbot Arena (LMSYS) Human preference-based ranking using Elo scores:

Leaderboard (example scores):
1. GPT-4o: 1290
2. Claude 3.5 Sonnet: 1271
3. Gemini 1.5 Pro: 1260
4. Llama 3.1 405B: 1250
...

Open LLM Leaderboard (HuggingFace) Automated benchmarks for open models:

HELM (Stanford) Holistic evaluation with many metrics:

Elo Rating System

def update_elo(winner_elo, loser_elo, k=32):
    expected_winner = 1 / (1 + 10 ** ((loser_elo - winner_elo) / 400))
    expected_loser = 1 - expected_winner

    new_winner_elo = winner_elo + k * (1 - expected_winner)
    new_loser_elo = loser_elo + k * (0 - expected_loser)

    return new_winner_elo, new_loser_elo

Interpreting Leaderboards

Elo DifferenceWin Probability
050%
10064%
20076%
40091%

Leaderboard Limitations

IssueMitigation
Selection biasRandom sampling
Prompt diversityTopic stratification
Position biasRandomize A/B order
Length biasEvaluate conciseness
TimeRatings change over time

Domain-Specific Leaderboards

DomainLeaderboard
CodingSWE-bench, LiveCodeBench
MathMATH leaderboard
SafetyHarmBench
RAGMTEB embeddings
AgentsAgentBench

Best Practices

leaderboardarenaelo

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.