Tech & AI · Live telemetry
LLM Knowledge & Coding Benchmarks
Published MMLU and HumanEval scores for leading large language models, as reported by each vendor. Not independently re-measured.
Embed
Embed this chart page anywhere with an iframe.
<iframe src="https://axiostats.com/datasets/llm-leaderboard/" title="AxioStats — LLM Knowledge & Coding Benchmarks" width="100%" height="480" loading="lazy"></iframe>
API endpoint
The same data as this page, served as static JSON — no key required.
curl https://axiostats.com/api/datasets/llm-leaderboard.jsonhttps://axiostats.com/api/datasets/llm-leaderboard.json
Executive summary
LLM Knowledge & Coding Benchmarks: GPT-4o tops the table at 88.7 MMLU, ahead of Claude 3.5 Sonnet (88.7). The spread from first to last (Mistral Large 2, 84) is 4.7 points.
Scores are as published by each vendor in their reports (not re-measured). MMLU measures knowledge accuracy; HumanEval measures code generation. See the 'source' column for the exact reference.
Data
| Model | MMLU | HumanEval | Source |
|---|---|---|---|
| GPT-4o | 88.7 | 90.2 | OpenAI GPT-4o report |
| Claude 3.5 Sonnet | 88.7 | 92 | Anthropic announcement |
| Llama 3.1 405B | 88.6 | 89 | Meta Llama 3.1 blog |
| DeepSeek-V3 | 88.5 | 82.6 | DeepSeek-V3 report |
| Gemini 1.5 Pro | 85.9 | 84.1 | Gemini 1.5 report |
| Mistral Large 2 | 84 | 81.1 | Mistral announcement |