← All matchups

Claude CodevsKimi CLI

3 prompts where both agents ran the exact same instructions. Each row is one prompt; the figures are whatever was measured or reported for that run.

This page holds the harness constant, not the model: the program driving a model changes the result as much as the model does, so the two are compared separately.

Claude Code and Kimi CLI ran the same prompt on 3 tasks, side by side. Cost, duration and outcome for each — 6 measured locally, no aggregate score.

Shared prompts
3
Compared
coding agents
Measured runs
6

No winner is declared. A measured run and a figure someone posted are not the same evidence, so they are never averaged into a ranking.

Vision Reveal — Creative Studio Hero
WebsitesReplayable pack
Claude Code
Fable 5$1.814m 25sCompletedmeasured
Kimi CLI
Kimi K3$0.194m 54sCompletedmeasured
Drift District — three-lap arcade racer
GamesReplayable pack
Claude Code
Kimi K3$0.7526m 55sFailedmeasured
Kimi CLI
Kimi K3$1.801h 32mTimed outmeasured
Warcraft 3 — mini RTS
GamesReplayable pack
Claude Code
Fable 5$18.5not measuredCompletedmeasured
Kimi CLI
Kimi K3$4.81not measuredCompletedmeasured

What this page shows, and how to read it

Claude Code is Anthropic's agentic coding CLI, and Kimi CLI is Moonshot AI's agentic coding CLI. A harness is the program that receives the prompt and works the task — editing files, running commands — while a model decides each step. On the prompts below, Claude Code drove Fable 5 and Kimi K3, and Kimi CLI drove Kimi K3.

The two sides share 3 prompts on this site: "Vision Reveal — Creative Studio Hero", "Drift District — three-lap arcade racer" and "Warcraft 3 — mini RTS". Across them, Claude Code completed 2 of 3 runs — of the others, 1 failed, and Kimi CLI completed 2 of 3 runs — of the others, 1 timed out. Claude Code has measured costs from $0.75 to $18.5 and measured durations from 4m 25s to 26m 55s. Kimi CLI has measured costs from $0.19 to $4.81 and measured durations from 4m 54s to 1h 32m.

How to read the outcomes. Completed means the prompt alone produced a working artefact. Failed means this run did not — the build broke, the result would not run, or the agent stopped short of a working state. That is a fact about one run on one prompt, not a verdict on the tool: the same programs complete other prompts elsewhere on this site. Timed out means the run was cut at its time limit with the artefact unfinished, so whatever it cost bought a partial result. Stopped and Interrupted mark runs ended from the outside before they concluded. A low cost attached to a run that did not finish is not a saving — it is the price of an attempt, which is why every figure on this page travels with its status.

Every figure above carries a trust label. Measured means SamePrompt ran it locally in Bench Arena, reconciled the tokens on the harness log and recomputed the cost from them. Reported means the author of the source announced the figure; it is shown as stated and cannot be verified here. The two are never summed, averaged or ranked, and this page declares no winner: a cheap run that failed and an expensive run that completed are two facts, not a score.