Monthly report
August 2026
48.1 / 100
+2.2 from July
Confidence: Medium
August continued the upward trend, led by coding and agents, while stronger evaluation practices added a modest reliability signal.
Category Scores
- Agents53
- Coding60
- Reasoning47
- Science35
- Multimodal49
- Robotics27
- Reliability38
Top signals this month
- Long-horizon agentic workflows became more concrete.
xAI released Grok 4.6 for coding, agentic tasks, and knowledge work, alongside Grok Bot for persistent cloud-based agent workflows.
- Gemini 3.7 Flash improved coding performance.
Google reported sizable gains over Gemini 3.6 Flash on debugging, issue resolution, and production-code evaluations.
- Science and evaluation quality both moved forward.
WeatherNext reported state-of-the-art cyclone forecasting results, while DeepMind piloted double-blind frontier-model evaluations designed to reduce benchmark contamination.