Agent Platform Scorecard

Who is building the agent operating system?

Twelve dimensions across seven providers. Re-graded as announcements land, with the market's reaction shown alongside.

8

OpenAI · Agent Capability

Can it actually do work? · updated 2026-06-01

Agent mode and Operator-style browsing ship real multi-step work, not just demos.

Field average: 6.6Dimension leader: Anthropic (9)
What this measures

Can the system take a goal and complete multi-step work autonomously (browse, write code, operate tools, recover from errors), not just answer questions? Measures shipped agentic behavior, not demos.

Across the web
Agent Capability across the field