Minecraft — five models build it from a two-line prompt
Reported by the author — we did not run this. What that means
Output





The prompt that built Minecraft
Make minecraft the game as a test. Do not stop until you are happy with the game and happy for me to review it. Use any tools you want to build the game. WEBGL, 3d stuff. Create your own assets if you want. Make it look great and feel like the real game
Runs
| Model | Harness | Status | Cost | Duration | Output | Trust | Artifact |
|---|---|---|---|---|---|---|---|
| Luna 5.6 (max reasoning)luna-5-6 | pi | Completed | $0.35 | not measured | not measured | Reported | — |
| DeepSeek V4 Flash-0731deepseek-v4-flash-0731 | pi | Completed | $0.25 | not measured | not measured | Reported | — |
| GLM 5.2 (high)glm-5-2 | pi | Completed | $0.30 | not measured | not measured | Reported | — |
| Kimi K3 (high)kimi-k3 | pi | Completed | $2.87 | not measured | not measured | Reported | — |
| Grok 4.5 (high)grok-4-5 | pi | Completed | $1.28 | not measured | not measured | Reported | — |
What this is, and what it proves
Two lines, no spec, no stack imposed: make Minecraft as a test, WebGL and 3D encouraged, and don't stop until it's ready for review. Five models took the same prompt through the same harness, over two rounds posted a day apart. Round one: Luna 5.6 at max reasoning produced the build the author called the clear winner — 343K tokens, $0.35 declared — against DeepSeek V4 Flash-0731's cheaper but weaker 366K-token, $0.25 attempt. Round two added three more at high effort: GLM 5.2 ($0.30, declared disappointing), Kimi K3 ($2.87, pretty good — it even added pigs), and Grok 4.5 ($1.28), which the author crowned the new overall winner, faster and better than Luna. Every figure on this page is the author's own declaration, nothing was measured here.
The prompt behind "Minecraft — five models build it from a two-line prompt" is published in full, and was run by 5 models: Luna 5.6 (max reasoning), DeepSeek V4 Flash-0731, GLM 5.2 (high), Kimi K3 (high) and Grok 4.5 (high). 5 of 5 finished, on figures reported by the author.
- Models
- 5
- Completed
- 5 of 5
- Cheapest completed run
- DeepSeek V4 Flash-0731 — $0.25
Figures reported by Matt. Not independently verified.
Reported, not measured
Two threads posted by the author on 1 and 2 August 2026; each clip here is the first eight seconds of his own captures. The prompt is published verbatim in a reply to the first thread and reproduced in full. Every figure is declared by the author in the threads: costs for all five runs, token totals only for round one (343K for Luna, 366K for DeepSeek — input and output combined, which is why they are not in the output-token fields). No durations were published. A follow-up reply states that all five models ran through the Pi harness (@pidotdev, pi.dev). The Luna clip's tweet is captioned 'GPT 5.6 Luna on Max reasoning' — the author attributes Luna to OpenAI elsewhere on his timeline. His verdicts, quoted: Luna 'absolutely crushed it', GLM 5.2 'disappointing', Kimi K3 'pretty good, and it even added pigs', Grok 4.5 'the new winner. Faster and even better than Luna, but more expensive'. The second thread is x.com/codermatt/status/2083711453010481193. Nothing on this page was measured by SamePrompt.