Twelve dimensions across seven providers. Re-graded as announcements land, with the market's reaction shown alongside.
Raw capability · updated 2026-06-01
Grok 4 closed most of the gap to the frontier at remarkable speed.
How smart is the assistant at the core reasoning, knowledge, and generation tasks people actually throw at it? Frontier benchmark standing, but weighted toward real assistant quality rather than leaderboard maxing.
Record reasoning/coding scores undercut by erratic, X-influenced behavior and protocol-skipping model changes.
Independent benchmarking that placed Grok 4 atop the Intelligence Index while flagging premium token pricing.