Coding comparison: Sol, Sonnet and Opus
One task, built three times from the same brief: the shot schedule rule and the client's shot ledger (S7 plan § 1.6), with their own tests and planted-bug proofs. 9 October 2026.
All three wrote correct code. Each passed the hidden suite, 67 of 67, and every build. They differ in how well their own tests guard the code, and in how closely the code reads like the rest of the project.
Opus came out best overall: first in the blind review, all six reference planted bugs caught, and code that matches the project's style. Sol wrote the leanest code, but the thinnest tests: 26 cases, and one of them crashed the whole run instead of failing. Sonnet wrote the most tests (128) and caught every planted bug, but added a rounding rule the plan doesn't have, which cost it in the blind review.
Scores
| Sol (GPT-6.1) | Sonnet 5.5 | Opus 5.5 | |
|---|---|---|---|
| Hidden acceptance suite (kept back from all three) | 67 / 67 | 67 / 67 | 67 / 67 |
| All three game builds (editor, server, client) | pass | pass | pass |
| Own tests written | 26 | 128 | 59 |
| Own planted bugs, each caught by its test | 6 / 6 | 14 / 14 | 11 / 11 |
| Six reference planted bugs: caught by their tests | 5 / 6 (one crashed the run) | 6 / 6 | 6 / 6 |
| Their tests run on the reference code | 23 / 26 | 124 / 128 | 57 / 59 |
| Blind review (scored by Sol, not told who wrote what) | 2nd: 10 / 9 / 9 / 8 | 3rd: 7 / 7 / 7 / 8 | 1st: 10 / 9 / 8 / 9 |
| Code written (the two files) | 197 lines | 345 lines | 305 lines |
| Wall time, including waits for the shared build lock | about 2 h 30 min | 2 h 12 min | 1 h 46 min (about an hour of it waiting) |
| Tokens | 17.8 M read (17.5 M cached), 135 k written; about US$3.60 at list price | 417 k (as the harness reports it) | 385 k (as the harness reports it) |
Blind review scores are correctness / robustness / code quality / tests, out of 10. The Claude token counts come from the agent harness, which does not split cached reads from writing, so they are not comparable in money with Sol's.
What separated them
- Sol's tests. One ledger test checks the list has four entries, then reads entry three without stopping when the check fails. Under the reference planted bug that empties the list, that read crashed the editor, and every test after it in the run went unrecorded. By our rule a crash never counts as caught. Without the crash, its check would have failed and caught the bug.
- Sonnet's extra rule. It wrote each boundary test both ways round (
E − S < CandE < S + C) to make the boundary exact whichever way a double sum rounds. The blind review found valid settings where a shot inside the carry still restarts the schedule, one rounding step off the plan's formula. Client and server share the code, so they would never disagree, but it is a rule the plan doesn't have. I rate it P2, not the review's P1. - Style. The project's code explains itself in comments. Opus and Sonnet matched that; Sol's two files carry almost none. Opus also gave validation one shared function, used by both the data check and the factory.
- Sol's retirement rule. Sol retires an answered send even above the acknowledgement marker; the header says only sends at or below it. The case can't arise from a correct server, so it is not a live defect.
- Where all three disagree with the reference. Each contestant's tests failed on the reference code on the same point: shot numbers. All three restart the next number at 1 as soon as a new generation is known, and emit the last possible number once. The reference resets only at the next shot, and stops one number short. Neither difference changes a shot in play, but the plan should state which is meant. It goes to the plan's owner.
How it was run
- Each built in its own copy of the project, with the real code and tests removed; the briefs were identical but for the name. Their logs show no look at the removed code or the project's history.
- The hidden suite is the project's own tests for these two files (Cadence, 50 cases; ShotLedger, 15; ShotHarness, 2), run against each version after it finished.
- The reference planted bugs are six of the project's own mutants of the real code. Each was applied to the real code and each contestant's tests run against it: a bug counts as caught only when a test that passed on the clean code fails, never when the run crashes.
- The blind review: GPT-6.1 Sol at medium effort, given the three versions as A, B and C, with no access to anything else.
- One task, one run each: a signal, not a verdict on the models.