Head to head: Kimi K3 vs xAI: Grok 4.5

This was a thin matchup on paper, but Kimi K3 still comes out ahead on the numbers. The catch: the only judged task showed enough order sensitivity that this result reads as a lean, not a rout.

By · Published · Updated

abstract symbolic representation of the story's core idea (editorial illustration in the spirit of New Yorker or The Atlantic)

Kimi K3 takes this head-to-head on the aggregate, 8.6 to 4.2, and the statistical read backs that up with a 75% confidence lean. That is enough to put it in front, but not enough to pretend this was a demolition. With only one scored task and no ties, the margin looks cleaner than the evidence really is.

Reader comments

Conversation for this story loads after sign-in.