AI Intelligence · Benchmarks

AI Benchmark Center

Live leaderboards across reasoning, coding, math, expert knowledge, and agent performance.

Current Leaders by Category

💻
Coding
CoreDev X
97.5
🛠
SWE-bench
CoreDev X
63.8
🧠
Reasoning
Athena Logic Pro
96.5
🔬
Science
Athena Logic Pro
92.1

Overall Model Ranking

(avg. normalized score across all benchmarks)
🥇NEW
CoreDev X
Axiom Dynamics
100.0
avg. score
🥈NEW
Athena Logic Pro
WisdomWorks
100.0
avg. score
🥉NEW
Aurora Prime v2
NovaSense AI
99.7
avg. score
4NEW
Apex Intellect
Veridian Labs
99.5
avg. score
5NEW
Sentient Core X
CogniPath
99.2
avg. score
6NEW
OmniVision AGI
Synthetica Inc.
98.9
avg. score
7NEW
TerraNova AGI
GlobalMind Inc.
98.8
avg. score
8NEW
ProtoGenius Code v1.0
SyntheDev
98.8
avg. score
💻

Coding

HumanEval · % pass@1

164 programming challenges requiring correct code from docstrings.

🥇
CoreDev XAxiom DynamicsNEW

Aug 5, 2026

97.5
🥈
ProtoGenius Code v1.0SyntheDevNEW

Aug 1, 2026

97.2
🥉
Genesis Code v3Origin LabsNEW

Aug 2, 2026

97
4
OmniCraft Coder v2.1AetherMindNEW

Jul 29, 2026

96.8
5
Arcana EngineerMysticTechNEW

Aug 4, 2026

96.3
🛠

SWE-bench

SWE-bench · % resolved

Real GitHub issues from popular open-source repositories.

🥇
CoreDev XAxiom DynamicsNEW

Aug 5, 2026

63.8
🥈
ProtoGenius Code v1.0SyntheDevNEW

Aug 1, 2026

62.5
🥉
Genesis Code v3Origin LabsNEW

Aug 2, 2026

61.5
4
OmniCraft Coder v2.1AetherMindNEW

Jul 29, 2026

60.1
5
Arcana EngineerMysticTechNEW

Aug 4, 2026

59.5
🧠

Reasoning

MMLU · % accuracy

57-subject knowledge test across STEM, humanities, and social sciences.

🥇
Athena Logic ProWisdomWorksNEW

Aug 7, 2026

96.5
🥈
Aurora Prime v2NovaSense AINEW

Aug 18, 2026

96.3
🥉
Apex IntellectVeridian LabsNEW

Aug 1, 2026

96.1
4
Sentient Core XCogniPathNEW

Aug 25, 2026

96
5
Sentinel AGI 1.0LuminaTechNEW

Aug 5, 2026

95.8
🔬

Science

GPQA · % accuracy

PhD-level science questions in biology, chemistry, and physics.

🥇
Athena Logic ProWisdomWorksNEW

Aug 7, 2026

92.1
🥈
Aurora Prime v2NovaSense AINEW

Aug 18, 2026

91.8
🥉
Apex IntellectVeridian LabsNEW

Aug 1, 2026

91.5
4
Sentient Core XCogniPathNEW

Aug 25, 2026

91.1
5
TerraNova AGIGlobalMind Inc.NEW

Aug 9, 2026

90.9

About These Benchmarks

MMLU (Massive Multitask Language Understanding)

Tests knowledge across 57 subjects including STEM, humanities, and social sciences. 14,000+ questions.

HumanEval (Coding)

164 hand-crafted programming challenges. Measures ability to produce correct code from docstrings.

MATH

12,500 competition math problems from AMC, AIME, and AMC 10/12. Tests advanced mathematical reasoning.

SWE-bench Verified

Real GitHub issues from popular open-source repos. Measures end-to-end software engineering capability.