← All matchupsMinecraft — five models build it from a two-line prompt
Grok 4.5 (high)vsKimi K3 (high)
1 prompt where both models ran the exact same instructions. Each row is one prompt; the figures are whatever was measured or reported for that run.
Grok 4.5 (high) and Kimi K3 (high) ran the same prompt on 1 task, side by side. Cost, duration and outcome for each — 0 measured locally, no aggregate score.
- Shared prompts
- 1
- Compared
- models
- Measured runs
- 0
No winner is declared. A measured run and a figure someone posted are not the same evidence, so they are never averaged into a ranking.
GamesReferenced
Grok 4.5 (high)
$1.28not measuredCompletedreported
Kimi K3 (high)
$2.87not measuredCompletedreported