XAI Grok 4 Benchmarks are showing it is the leading model. On Humanity’s Last Exam in had scores of 35 and 45 for reasoning is a big improvement from about 21 for other top models.
If these leaked Grok 4 benchmarks are correct, 95 AIME, 88 GPQA, 75 SWE-bench, then XAI has the most powerful model on the market.
The GPQA for Grok and SWE Bench rankings for Grok 4 code will also top the rankings.