Read the issue and the code.
Repository-level tasks test more than isolated code generation: the agent must locate the relevant behavior and understand its context.
Quality report / SWE-bench Verified
Amber scored higher than GPT-5.6 Sol in an evaluation of 50 tasks from SWE-bench Verified: 88% vs 86%. That’s 44 real repository issues resolved, compared with Sol’s 43.
Internal evaluation · Fixed 50-task subset · Not a full SWE-bench Verified score
01 / Matched evaluation
Issues resolved on the same 50-task SWE-bench Verified subset. Higher is better.
88%44 / 50 resolved
86%43 / 50 resolved
Amber leads by two percentage points. One more issue resolved in this evaluation; not evidence of overall superiority.
Amber run: 2 September 2026. Sol run: 28 August 2026. Both submitted all 50 patches; neither run recorded an evaluation error.
02 / What quality means here
SWE-bench uses real issues from open-source repositories. The agent must work with an existing codebase and produce a patch. The evaluation checks whether that patch resolves the issue without breaking the required regression tests.
Repository-level tasks test more than isolated code generation: the agent must locate the relevant behavior and understand its context.
The model operates through a coding-agent scaffold. Results measure the model, its instructions and its tools together.
“Resolved” is the benchmark harness result, not a subjective rating of how polished the model’s explanation sounds.
03 / Experimental setup
We ran both configurations against the same fixed subset. The comparison is between these specific agent configurations, not every possible setting of either model.
xhigh), self-hosted by 2BA.gpt-5.6-sol, medium reasoning, accessed through the Responses API. Model name as recorded in the benchmark.04 / Interpretation & limitations
These are internal results on 50 of the 500 Verified tasks, not an official full-suite submission. A one-task difference does not establish statistical equivalence or superiority.
Amber’s extra-high reasoning configuration was selected after several prompt and effort experiments on this same subset. It is not an untouched holdout, and that tuning can favor the reported result.
Both configurations resolved 38 of the same tasks. Amber alone resolved six; Sol alone resolved five. Similar scores do not mean interchangeable answers on every issue.
Single runs do not quantify run-to-run variation. Public benchmark exposure may also affect results. Validate the model on your own repositories and review generated patches.
The headline uses the completed NVFP4 Amber reference run and the matched Sol run in our benchmark scorecard. Later hardware and incomplete experimental runs are not pooled into this comparison.
05 / Your code is the next test
Try 2BA with your coding agent. Pair your browser, create an API key and configure your tool.