Coding comparison: Sol, Sonnet and Opus

One task, built three times from the same brief: the shot schedule rule and the client's shot ledger (S7 plan § 1.6), with their own tests and planted-bug proofs. 9 October 2026.

All three wrote correct code. Each passed the hidden suite, 67 of 67, and every build. They differ in how well their own tests guard the code, and in how closely the code reads like the rest of the project.

Opus came out best overall: first in the blind review, all six reference planted bugs caught, and code that matches the project's style. Sol wrote the leanest code, but the thinnest tests: 26 cases, and one of them crashed the whole run instead of failing. Sonnet wrote the most tests (128) and caught every planted bug, but added a rounding rule the plan doesn't have, which cost it in the blind review.

Scores

Sol (GPT-6.1)Sonnet 5.5Opus 5.5
Hidden acceptance suite (kept back from all three)67 / 6767 / 6767 / 67
All three game builds (editor, server, client)passpasspass
Own tests written2612859
Own planted bugs, each caught by its test6 / 614 / 1411 / 11
Six reference planted bugs: caught by their tests5 / 6 (one crashed the run)6 / 66 / 6
Their tests run on the reference code23 / 26124 / 12857 / 59
Blind review (scored by Sol, not told who wrote what)2nd: 10 / 9 / 9 / 83rd: 7 / 7 / 7 / 81st: 10 / 9 / 8 / 9
Code written (the two files)197 lines345 lines305 lines
Wall time, including waits for the shared build lockabout 2 h 30 min2 h 12 min1 h 46 min (about an hour of it waiting)
Tokens17.8 M read (17.5 M cached), 135 k written; about US$3.60 at list price417 k (as the harness reports it)385 k (as the harness reports it)

Blind review scores are correctness / robustness / code quality / tests, out of 10. The Claude token counts come from the agent harness, which does not split cached reads from writing, so they are not comparable in money with Sol's.

What separated them

How it was run