Agentic
★

Benchmarks

4 benchmarks · updated 1h ago · Methodology →
GAIAgeneral
τ-benchtool use
SWE-benchcode
WebArenaweb
Generalist assistant tasks — web browsing, code execution, multimodal reasoning. Higher is better, 0–100.

Biggest movers

Last 7 days →
OP
Operator 2.0
71.2
▲ 3.4
A7
Agent-72B
68.9
▲ 1.1
SM
SmolAgents
61.3
NEW

Leaderboard42

Full table →
1
OP
Operator 2.0Proprietary · web
71.2▲ 3.4
2
A7
Agent-72BOpen weights · Apache 2.0
68.9▲ 1.1
3
LG
LangGraph + ClaudeFramework · proprietary model
66.4▼ 0.6
4
CA
CrewAI + GPT-5Framework · proprietary model
64.1▲ 0.8
5
AG
Agno + Agent-72BFramework · open weights
62.7▲ 2.2
Showing 1 – 5 out of 42
‹123…9›

How scores work. Normalized 0–100 from public eval runs, re-run hourly. Deltas compare against 7 days ago. Open-weights entries are marked so you know what you can actually run.