How we got to the one-sided benchmark
I started with a simple, slightly mischievous question: can we find a small task on which Sol is dramatically better than Opus? The dangerous version of that question is a search for a winning screenshot. The useful version is a search for a result that survives hostile review.
That distinction shaped the entire project. The comparison in this article is GPT-5.6 Sol versus Claude Opus 5, with DeepSeek V4 Pro as a third comparator. Earlier models and experiments are part of the trail, not part of the final claim.
The first idea was too broad
The first wave used agentic software tasks: C++ coroutine ownership, Go POSIX process supervision, Python coordinate transforms, Python codemods, and Rust bounded joins. They sounded like the right kind of “real work.” They were also difficult to interpret.
The agentic runs produced weak or fail-open evidence. The coordinate family had a verifier contradiction: the task required output.py, while the verifier rejected the presence of that file. Every model was rejected for the same harness reason. Two Sol POSIX runs implemented plausible repairs but killed the persistent shell with exit $rc, so Harbor could not collect the resulting artifact. A DeepSeek codemod run hit the runaway guard. Those rows remain in the raw record, but they cannot honestly be called clean capability failures.
The clean agentic remainder was essentially a ceiling: Sol 15/15, Opus 17/17, and DeepSeek 16/16 after the invalid family was excluded. There was no semantic Sol-versus-Opus gap. The lesson was not that agent benchmarks are useless. It was that a task can be “real” and still be a poor measuring instrument when the agent shell, verifier, artifact collection, and model behavior are entangled.
The async experiments made the problem even clearer. An async bounded-map task could make one model look better, then another task could reverse the direction. The evidence flipped by task. That is exactly what a broad leaderboard-style benchmark hides: a global average can conceal several incompatible failure modes. We kept those tasks as history and stopped treating them as a route to a one-sided result.
The CMake temptation, and the deadline we abandoned
The next attractive result came from CMake/linker repairs. In discovery, all three models eventually solved the examples, but Sol was faster: approximately 40, 52, and 73 seconds, versus Opus at approximately 136, 186, and 242 seconds. We designed a new holdout around a 120-second “natural” agent deadline and preregistered a rule requiring Sol to succeed on at least five of six tasks while Opus succeeded on at most one.
That was a coherent operational question, but it was not the question the user wanted answered. A time cutoff can turn “eventually solves the task” into “fails,” and the cutoff was selected after observing discovery runtimes. Even with the disclosure and frozen holdout, it felt like scoring latency as capability. The user rejected time cutoffs as cheating. We retired the CMake holdout from the one-sided claim rather than moving the goalposts around it.
That rejection was productive. It forced the benchmark toward tasks where the answer itself is the evidence.
Small gotchas became the right unit
We pivoted to non-agentic microtasks: subtitles, structured logs, CSV, citations, diagnostics, then smaller format and text gotchas. Each case supplies an operation and source text and asks for one JSON value. No tools, files, network, agent loop, or scoring deadline are involved. The verifier compares the returned value exactly.
The first direct-micro result looked promising: Sol scored 140/140 and Opus 127/140. But this still did not satisfy the project’s confirmed one-sided rubric. Opus was too accurate overall, and most of its 13 misses were serialization or wrapper errors rather than failures to understand the transformation. That was a useful result about exact protocol completion, not a universal reasoning claim.
Subtitles showed why audit mattered more than a large percentage gap. One exploratory subtitle family initially looked spectacular: Sol 1/9, Opus 7/9, DeepSeek 8/9. The prompt required a format field but never specified its literal spelling or case; the hidden oracle required lowercase srt or vtt. Sol transformed the cues correctly but returned SRT or WebVTT. The apparent Sol loss was a scoring-contract artifact. Once the ambiguous field was removed in a fresh follow-up, Sol and Opus both scored 12/12. We retired the original subtitle family for capability claims and preserved its raw rows.
That was one of the most important findings in the whole journey: a repeated hidden convention can manufacture a beautiful, statistically significant, completely misleading one-sided benchmark. The benchmark must test the model, not the evaluator’s private preference.
Unicode supplied the mechanism
Among the gotcha packs, Unicode was the first mechanism that survived direct inspection. The task was to perform a local operation while preserving every unrelated scalar. Opus often normalized decomposed text—a plus combining acute became á—or converted CRLF to LF. Sol preserved the original code points. A fresh Unicode-v2 follow-up made the contract explicit with a code-point ledger, prohibited normalization and line-ending conversion, and checked grapheme segmentation independently with Python regex and Swift Character.
This was not a broad claim that one model “understands Unicode” and the other does not. The unseen follow-up was 15/15 for Sol, 11/15 for Opus, and 13/15 for DeepSeek. Opus’s four misses preserved grapheme boundaries but changed the underlying representation. The mechanism was real, the cases were related, and the result remained exploratory.
ANSI control strings then exposed a second mechanism. Models had to remove one precisely defined control sequence while preserving ordinary bracket-like text, C1 controls, payload, combining marks, and line endings. Opus sometimes left control payload behind, deleted controls that were supposed to remain, or normalized unrelated text. Sol was consistently lossless.
The first ANSI follow-up also exposed an endpoint problem: Opus produced null content on control-heavy probes. Every Opus response that contained valid content was correct, so that run showed an end-to-end output-availability gap, not necessarily a parser gap. We therefore froze a fresh, stricter ANSI replication with unique structures and retained every transport outcome instead of silently converting missing output into semantic failure.
M2: the first clean confirmation
M2 separated discovery from inference. The ANSI direction was selected after inspecting discovery. Then an untouched 24-case holdout was opened under a frozen lock. The models never received expected answers or case identifiers that would reveal the scoring key. All Sol and Opus answers were schema-valid, and the frozen ANSI oracle independently recomputed the expected values. DeepSeek had one invalid or malformed holdout response.
The holdout result was:
| Model | Correct |
|---|---|
| Sol | 22/24 |
| Opus | 4/24 |
| DeepSeek V4 Pro | 15/24 |
There were 18 Sol-only cases and zero Opus-only cases. The exact paired two-sided McNemar p-value was 7.629e-6. This is the strongest formal result in the project. It is a result about exact, lossless Unicode and ANSI/control surgery on the frozen holdout—not about general intelligence or overall model quality.
R1: trying to make the claim harder to keep
We then attempted a stricter fresh replication, R1, with 30 unique ANSI structures and a preregistered absolute Sol gate of 85%. The scores were Sol 25/30, Opus 0/30, and DeepSeek 8/30. The paired direction was 25–0 and the exact paired p-value was 5.96e-8.
But 25/30 is 83.3%, one case short of the frozen 26/30 threshold. R1 therefore formally failed. We did not lower the threshold after seeing the result. The honest status is no_confirmed_gap for R1, despite the huge observed separation. That failed confirmation is more valuable than a retrofitted success because it shows the rule was real.
The larger characterization and reservoir runs showed the same direction—88/96 versus 20/96 in the lossless characterization, and 147/236 versus 13/236 in the ordered reservoir prefix—but they were descriptive. They contained related variants, and the reservoir stopped when provider quota became the binding limit. The additional project budget still had room, but the provider did not. We stopped rather than manufacture failures or pretend the remaining money was usable.
AlphaOx read the work skeptically
At the end, I asked OpenCode Go’s free Ox Alpha to audit the construction, tests, results, and article from a disposable snapshot. AlphaOx confirmed that the analyzer and targeted tests were reproducible, and it agreed with the M2 and R1 arithmetic. It also found the caveats that belong in the article:
- the two ANSI scanners share a hand-written grammar, so the evidence establishes correctness relative to that explicit grammar, not all of ECMA-48;
- the cases are adversarial fixed fixtures with recurring lexical mechanisms, not a random sample of text;
- provider routing differed across endpoints, especially for Opus and DeepSeek;
- DeepSeek produced 52 of 58 output failures in the relevant R1 analysis, including 49 length stops at the 8,192-token cap;
- the R1 lock contains a confusing
source_base_headname even though the exposure ledger binds the actual source commit used for the run.
Those caveats do not erase the M2 result. They define its perimeter.
I also asked AlphaOx to inspect the benchmark questions and answer keys directly. It manually reviewed all 24 published M2 holdout pairs and judged every key consistent with the stated operation. It then reviewed 18 of the 30 R1 headline pairs, again finding no incorrect key. The free endpoint repeatedly ended longer sessions with an unknown completion state, so I am not presenting that as a complete external review of R1: 12 R1 headline pairs, the 96 characterization cases, and the reservoir were not manually checked pair by pair by AlphaOx. They remain covered by the frozen generators, scanner agreement, and local tests, but that is weaker than a fully independent manual oracle audit.
The claim we can defend
The final claim is deliberately narrow:
On exact lossless Unicode and ANSI/control surgery, Sol was dramatically more reliable than Opus in an untouched M2 holdout. The same direction appeared in a larger fresh R1 run, but R1 missed its own absolute Sol-accuracy gate by one case and is not a successful formal replication.
We cannot say that Sol is universally better than Opus. We cannot pool discovery, holdout, characterization, and quota-stopped reservoir rows into one triumphant percentage. We cannot call the subtitle artifact a model weakness. We cannot call null output, provider routing, or a killed shell a clean semantic failure. And we cannot turn a 120-second deadline into a general capability claim.
Lessons and reproduction
The journey changed my definition of a good benchmark. The best one-sided benchmark is not the one with the biggest first gap. It is the one that survives attempts to explain the gap away: explicit contracts, untouched holdouts, exact code-point scoring, deterministic independent checks, fail-closed analyzers, raw rows preserved, and thresholds that remain binding after the result arrives.
The complete result is published in The one-sided benchmark. The frozen source commit was 523a785456b3a981da56fc73f3e67433b3cd11d9.
The raw evidence, incident records, admissions, charge ledgers, and completion manifests are the record. Every reported score should be regenerated from those artifacts, not from this narrative.