Head to head: Kimi K3 vs xAI: Grok 4.5
This was a thin matchup on paper, but Kimi K3 still comes out ahead on the numbers. The catch: the only judged task showed enough order sensitivity that this result reads as a lean, not a rout.
By RuntimeWire · Published · Updated

Kimi K3 takes this head-to-head on the aggregate, 8.6 to 4.2, and the statistical read backs that up with a 75% confidence lean. That is enough to put it in front, but not enough to pretend this was a demolition. With only one scored task and no ties, the margin looks cleaner than the evidence really is.