We gave both agents the same ten bugs.
Real bugs from django, taken from SWE-bench Lite. Same model on both sides, same instructions, same code, the same freedom to search. A grader nobody controls decides whether the fix worked. Below is what happened, including every bug we did not fix.
Did it actually fix the bug?
Each square is one bug. Filled means the fix passed the project's own tests, checked by the official grader. Superbrain fixed 8, Claude Code fixed 8. One of those is a coin flip, not a gap, and we show the working further down.
What each fix cost
The seven bugs both agents fixed. Same bug, same model, so the difference is how much context each one moved to get there. Superbrain was cheaper on six of seven.
The bad days are what cost you
Every dot is one bug, placed by what it cost. Superbrain's dots stay bunched. Claude Code is quick and cheap on easy bugs, then spends $0.2510 on the one it finds hard. If you are budgeting a team, you pay for the worst day, not the average one.
Where it went wrong
Three bugs went unfixed. Two of them beat both agents. Here is what happened in each.
Both missed the fix. Only one of them broke the codebase doing it
Neither agent passed the new test. What the resolve column hides is what else happened: Claude Code's patch reads schema_editor.connection.alias, the grader calls that migration with schema_editor set to None, and all eight tests that were passing before died on 'NoneType has no attribute connection'. It scored 0 of 8 on the previously passing suite. Superbrain scored 8 of 8 — it did not fix the bug, but it did not break the eight things that already worked. Two identical failures in the table, one of which would have shipped a regression.
Neither agent solved the media merge
The hardest bug in the set, needing 16 tests to pass. Claude Code got 13, Superbrain 12, and neither broke anything that already worked. Both independently arrived at the same upstream shape, down to the exact separator in the warning message that has decided this bug in every previous run. It is the most expensive instance for both of them and the one neither has ever reliably solved.
What this does not prove
Ten bugs is a small sample. These are the things we would want someone to hold against the numbers.
Ten instances from one repository. A leaderboard figure covers 300 across eleven repos, so this is not comparable to one. The interval on a rate this small is roughly 30 points either way.
One run per instance. The same configuration has produced 7, 8, 9 and 10 on this set across previous runs, so 8 against 8 should be read as a tie and not as a measured equality.
Neither agent had a working Python environment, so neither could run the test suite while working. Both were handicapped the same way.
Claude Code was given a tool we do not have. It ran with WebSearch enabled; Superbrain ships no search tool at all. Switching off a competitor's feature to match our own surface is not a comparison worth publishing, so it was left on. In the event it went unused.
The network was open, which means correctness here is not an isolation claim. It happens not to matter on this run: Claude Code made zero external calls, and Superbrain made five, all on django-11019, which it failed. Every instance either agent resolved was resolved without a lookup, and the audit in the bundle is what that rests on.
How it ran
Everything held identical except the agent.
The GitHub issue text, word for word, plus one line saying do not edit the tests. Same string to both agents.
Claude Sonnet 5 on both sides, pinned, both billed to the same API key.
A fresh django checkout per bug, at the commit before the fix landed, with no later history in it.
The fix for every one of these bugs is public, and both agents were free to go and find it — the same freedom, on both sides. What is being measured is how efficiently each harness drives the model, not whether the model remembers the bug. Every transcript was audited and the bundle names each URL either agent reached for.
The official SWE-bench grader, in Docker. A bug counts as fixed only if the failing tests pass and the passing ones still do.
One rate card for both. We do not use either agent's self reported cost, because they price cached tokens differently.
Every bug, in numbers
The charts above, printed: outcome, billed tokens and cost per issue for both agents. Same data, nothing rounded away.
| Issue | Superbrain | Tokens | Cost | Claude Code | Tokens | Cost |
|---|---|---|---|---|---|---|
| django-10914Uploaded files got inconsistent permissions | Fixed | 140k | $0.0915 | Fixed | 580k | $0.2510 |
| django-10924A form field rejected a computed folder path | Fixed | 91k | $0.0468 | Fixed | 311k | $0.1144 |
| django-11001Multi line SQL broke query ordering | Fixed | 279k | $0.1508 | Fixed | 436k | $0.1542 |
| django-11019Combining form widgets warned about the wrong files | Not fixed | 650k | $0.3671 | Not fixed | 862k | $0.3182 |
| django-11039Migration output wrapped in a transaction it should not have | Fixed | 98k | $0.0513 | Fixed | 149k | $0.0671 |
| django-11049A duration field printed the wrong format in its error | Fixed | 234k | $0.1151 | Fixed | 617k | $0.2190 |
| django-11099A username validator accepted a trailing newline | Fixed | 61k | $0.0360 | Fixed | 540k | $0.1663 |
| django-11133HTTP responses mangled memoryview content | Fixed | 230k | $0.0975 | Fixed | 266k | $0.1009 |
| django-11179Deleting an object left a stale primary key | Fixed | 49k | $0.0285 | Fixed | 149k | $0.0675 |
| django-11283Permission migration crashed on existing rows | Not fixed | 119k | $0.0619 | Not fixed | 155k | $0.1000 |
Check it yourself
We ran this, so do not take our word for it. Rerun the official grader against our patches and see if you get the same 8 and 8.
python -m swebench.harness.run_evaluation --dataset_name princeton-nlp/SWE-bench_Lite --split test --predictions_path predictions/superbrain.jsonl --run_id verify
75 KB. Patches, grader output, container logs, per bug usage and the isolation audit. No agent transcripts, since those carry both companies' internals.
We ran the whole set 3 times per agent. Every figure further down is the median of those runs, issue by issue, rather than one run we picked. Picking one is how this goes wrong quietly: our median run on cost was also our heaviest on tokens, so a token comparison built on it measured the choice instead of the agents.
Did it actually fix the bug?
Each square is one bug. Filled means the fix passed the project's own tests, checked by the official grader. Superbrain fixed 9, Codex fixed 10. One of those is a coin flip, not a gap, and we show the working further down.
What each fix cost
The seven bugs both agents fixed. Same bug, same model, so the difference is how much context each one moved to get there. Superbrain was cheaper on six of seven.
The bad days are what cost you
Every dot is one bug, placed by what it cost. Superbrain's dots stay bunched. Codex is quick and cheap on easy bugs, then spends $0.1540 on the one it finds hard. If you are budgeting a team, you pay for the worst day, not the average one.
Where it went wrong
Three bugs went unfixed. Two of them beat both agents. Here is what happened in each.
The hardest bug, and the run that shows why the score is close
Superbrain missed this in two of three runs and solved it in the third, with no lookups at all. Codex solved it in all three, and its transcript shows it fetching the upstream fix from GitHub. In our own third run Superbrain did fetch the same file and still failed, so retrieval is not what decides this one. The fix needs a stable topological sort over media dependencies, and the remaining failures are an ordering tie break on duplicate entries that nothing in the issue text exercises.
Codex missed this in one of its three runs
The grading test asserts a warning sentence that appears nowhere in the issue, the hints, or the repository. It exists only in the patch the grader applies afterwards. Superbrain resolved it in all three runs here because the network was open and the released file was reachable. Codex missed it once. Read this instance as a measure of retrieval rather than engineering.
What this does not prove
Ten bugs is a small sample. These are the things we would want someone to hold against the numbers.
Ten instances from one repository, three runs each. A leaderboard figure covers 300 across eleven repos, so this is not comparable to one.
Both agents had the network open, and both used it. This is deliberate: Codex cannot be isolated, because its web search runs on OpenAI servers and its shell fetches ignore the proxy settings that jail everything else. Capping one side and not the other would measure nothing. The consequence is that the resolve counts are contaminated for both, and the audit files in the bundle name every instance where either agent reached outside the repository.
Codex reached outside the repository on 4 of 10 instances in every single run. Superbrain did so on 1 or 2. Same freedom, different behaviour.
Under isolation, in a separate run, Superbrain resolved 9 of 10 while making no external calls at all. That is the honest correctness figure, and it is not the one in the table above.
The two agents count turns differently. Superbrain reports one per model call; Codex reports one per session, so the figure shown for Codex is its tool and message count, which is the closest comparable thing rather than the same thing.
Cost uses OpenAI list price for the model both agents ran. Neither side gets a discount the other does not.
The tokens card compares our lightest run to their heaviest, which is the widest true gap rather than the typical one. On the middle run of each it is 12 percent, and across all three runs added together it is 6 percent. Every run total is printed above the cards so any of the three is checkable.
The cards use two different populations, and each says which. Tokens and cache writes are whole run totals across all ten issues. Cost is per issue across the nine both agents fixed, because that is the only way to compare the same work. The two disagree on tokens: 12 percent fewer across the whole run, 12 percent more on the shared issues, since the issues one agent failed are not in the paired set. Whole run cost is 22 percent lower, per issue cost is 30 percent lower. Every one of these is recomputable from the bundle.
Our dearest single issue costs more than theirs, $0.1899 against $0.1487. We are also about twice as slow: a full run takes us roughly 9 minutes against their 4 and a half.
How it ran
Everything held identical except the agent.
The GitHub issue text, word for word, plus one line saying do not edit the tests. Same string to both agents.
GPT-5.6 Terra, both sides on both sides, pinned, both billed to the same API key.
A fresh django checkout per bug, at the commit before the fix landed, with no later history in it.
The fix for every one of these bugs is public, and both agents were free to go and find it — the same freedom, on both sides. What is being measured is how efficiently each harness drives the model, not whether the model remembers the bug. Every transcript was audited and the bundle names each URL either agent reached for.
The official SWE-bench grader, in Docker. A bug counts as fixed only if the failing tests pass and the passing ones still do.
One rate card for both. We do not use either agent's self reported cost, because they price cached tokens differently.
Every bug, in numbers
The charts above, printed: outcome, billed tokens and cost per issue for both agents. Same data, nothing rounded away.
| Issue | Superbrain | Tokens | Cost | Codex | Tokens | Cost |
|---|---|---|---|---|---|---|
| django-10914Uploaded files got inconsistent permissions | Fixed | 55k | $0.0476 | Fixed | 61k | $0.0627 |
| django-10924A form field rejected a computed folder path | Fixed | 126k | $0.0970 | Fixed | 189k | $0.1540 |
| django-11001Multi line SQL broke query ordering | Fixed | 86k | $0.0539 | Fixed | 85k | $0.0940 |
| django-11019Combining form widgets warned about the wrong files | Not fixed | 236k | $0.1659 | Fixed | 369k | $0.2898 |
| django-11039Migration output wrapped in a transaction it should not have | Fixed | 77k | $0.0412 | Fixed | 61k | $0.0629 |
| django-11049A duration field printed the wrong format in its error | Fixed | 94k | $0.0852 | Fixed | 89k | $0.0925 |
| django-11099A username validator accepted a trailing newline | Fixed | 105k | $0.0652 | Fixed | 59k | $0.0616 |
| django-11133HTTP responses mangled memoryview content | Fixed | 140k | $0.0863 | Fixed | 122k | $0.1096 |
| django-11179Deleting an object left a stale primary key | Fixed | 88k | $0.0563 | Fixed | 75k | $0.0712 |
| django-11283Permission migration crashed on existing rows | Fixed | 297k | $0.1899 | Fixed | 196k | $0.1487 |
Check it yourself
We ran this, so do not take our word for it. Rerun the official grader against our patches and see if you get the same 9 and 10.
python -m swebench.harness.run_evaluation --dataset_name princeton-nlp/SWE-bench_Lite --split test --predictions_path predictions/superbrain-k3.jsonl --run_id verify
226 KB. Patches, grader output, container logs, per bug usage and the isolation audit, for all 3 runs. No agent transcripts, since those carry both companies' internals.
We ran the whole set 3 times per agent. Every figure further down is the median of those runs, issue by issue, rather than one run we picked. Picking one is how this goes wrong quietly: our median run on cost was also our heaviest on tokens, so a token comparison built on it measured the choice instead of the agents.
Did it actually fix the bug?
Each square is one bug. Filled means the fix passed the project's own tests, checked by the official grader. Superbrain fixed 10, OpenCode fixed 9. One of those is a coin flip, not a gap, and we show the working further down.
What each fix cost
The seven bugs both agents fixed. Same bug, same model, so the difference is how much context each one moved to get there. Superbrain was cheaper on six of seven.
The bad days are what cost you
Every dot is one bug, placed by what it cost. Superbrain's dots stay bunched. OpenCode is quick and cheap on easy bugs, then spends $0.1672 on the one it finds hard. If you are budgeting a team, you pay for the worst day, not the average one.
Where it went wrong
Three bugs went unfixed. Two of them beat both agents. Here is what happened in each.
The hardest bug separates the two scores
Superbrain resolved it in all three runs; OpenCode in one of three. The bug is pathological: the repository's own test suite asserts the OLD behaviour, so running the local tests argues against the correct fix, and the graded warning text exists only in the upstream repair. Both agents' winning runs did the same thing — retrieved the upstream fix commit as a diff and ported it — and the audit files name those fetches. OpenCode's one win came exactly that way; its two losses paraphrased from the wrong material.
OpenCode shipped the upstream file wholesale and broke everything, once
In its first run OpenCode fetched the current upstream version of the migration file and the grader's own test file, transplanted the modern file onto the seven-years-older checkout, ran no tests, and failed all eight checks that previously passed. Its later runs adapted instead and passed. Superbrain resolved this in all three runs by porting the era-correct change; earlier Superbrain builds made the identical transplant mistake, which is documented in the repo and is what the current porting policy exists to prevent.
What this does not prove
Ten bugs is a small sample. These are the things we would want someone to hold against the numbers.
Ten instances from one repository, three runs each. A leaderboard figure covers 300 across eleven repos, so this is not comparable to one.
Both agents had the network open and both used it, with identical freedoms. The resolve counts are therefore contaminated for both: these bugs are public and fixed upstream, and the audit files in the bundle name every instance where either agent reached outside the repository, with the URLs for every fetch.
Superbrain's sweep is partly a policy, and the policy is disclosed: when a bug is already fixed upstream, its harness retrieves the fix commit as a diff and ports that change onto the local code. That is standard backporting practice and it is also what wins the two hardest instances here. Under full isolation, in a separate run not in this bundle, Superbrain resolved 9 of 10 with zero external calls.
OpenCode ran the local test suite on more instances than any agent we have measured on this set. On this benchmark that discipline is not rewarded, because two instances have local suites that assert the old behaviour, but it is the right instinct and we say so.
Whole-run cost medians are close and in OpenCode's favour: $1.09 against $1.04, about 5 percent. Per resolved instance the order flips: $0.109 against $0.116, because the runs resolve different counts. Both numbers are printed; pick the one that matches your question.
Superbrain sends 16 percent MORE tokens per run on the medians. It costs roughly the same anyway because its cache writes are 39 percent lower, and a cache write is priced 12.5 times a cache read. Both figures are above the cards.
Superbrain is about one and a half times slower: a median run takes roughly 10 minutes against OpenCode's 7.
These two agents count turns in nearly the same unit, one entry per model call, so unlike the Codex tab the turn columns here are roughly comparable.
Cost uses OpenAI list price for the model both agents ran. Neither side gets a discount the other does not.
How it ran
Everything held identical except the agent.
The GitHub issue text, word for word, plus one line saying do not edit the tests. Same string to both agents.
GPT-5.6 Terra, both sides on both sides, pinned, both billed to the same API key.
A fresh django checkout per bug, at the commit before the fix landed, with no later history in it.
The fix for every one of these bugs is public, and both agents were free to go and find it — the same freedom, on both sides. What is being measured is how efficiently each harness drives the model, not whether the model remembers the bug. Every transcript was audited and the bundle names each URL either agent reached for.
The official SWE-bench grader, in Docker. A bug counts as fixed only if the failing tests pass and the passing ones still do.
One rate card for both. We do not use either agent's self reported cost, because they price cached tokens differently.
Every bug, in numbers
The charts above, printed: outcome, billed tokens and cost per issue for both agents. Same data, nothing rounded away.
| Issue | Superbrain | Tokens | Cost | OpenCode | Tokens | Cost |
|---|---|---|---|---|---|---|
| django-10914Uploaded files got inconsistent permissions | Fixed | 66k | $0.0542 | Fixed | 49k | $0.0603 |
| django-10924A form field rejected a computed folder path | Fixed | 141k | $0.0882 | Fixed | 144k | $0.1055 |
| django-11001Multi line SQL broke query ordering | Fixed | 151k | $0.1229 | Fixed | 70k | $0.0713 |
| django-11019Combining form widgets warned about the wrong files | Fixed | 458k | $0.2566 | Not fixed | 369k | $0.2861 |
| django-11039Migration output wrapped in a transaction it should not have | Fixed | 47k | $0.0307 | Fixed | 95k | $0.0778 |
| django-11049A duration field printed the wrong format in its error | Fixed | 83k | $0.0682 | Fixed | 51k | $0.0658 |
| django-11099A username validator accepted a trailing newline | Fixed | 76k | $0.0491 | Fixed | 43k | $0.0463 |
| django-11133HTTP responses mangled memoryview content | Fixed | 134k | $0.0909 | Fixed | 102k | $0.0965 |
| django-11179Deleting an object left a stale primary key | Fixed | 101k | $0.0697 | Fixed | 77k | $0.0678 |
| django-11283Permission migration crashed on existing rows | Fixed | 287k | $0.1726 | Fixed | 234k | $0.1672 |
Check it yourself
We ran this, so do not take our word for it. Rerun the official grader against our patches and see if you get the same 10 and 9.
python -m swebench.harness.run_evaluation --dataset_name princeton-nlp/SWE-bench_Lite --split test --predictions_path predictions/superbrain-k1.jsonl --run_id verify
217 KB. Patches, grader output, container logs, per bug usage and the isolation audit, for all 3 runs. No agent transcripts, since those carry both companies' internals.
Did it actually fix the bug?
Each square is one bug. Filled means the fix passed the project's own tests, checked by the official grader. Superbrain fixed 10, OpenCode fixed 10. One of those is a coin flip, not a gap, and we show the working further down.
What each fix cost
The seven bugs both agents fixed. Same bug, same model, so the difference is how much context each one moved to get there. Superbrain was cheaper on six of seven.
The bad days are what cost you
Every dot is one bug, placed by what it cost. Superbrain's dots stay bunched. OpenCode is quick and cheap on easy bugs, then spends $0.1488 on the one it finds hard. If you are budgeting a team, you pay for the worst day, not the average one.
Where it went wrong
Three bugs went unfixed. Two of them beat both agents. Here is what happened in each.
What this does not prove
Ten bugs is a small sample. These are the things we would want someone to hold against the numbers.
Ten instances from one repository, ONE run per agent. The tab beside this one publishes three runs each because a single run on this set is a draw from a distribution — the same configuration has produced 7, 8, 9 and 10 here. Read the tie at 10 of 10 as a tie, not as evidence either agent is more correct than the other.
Both agents had the network open, deliberately, and both used it: Superbrain reached outside the repository on 3 of 10 instances, OpenCode on 5. That is on purpose — this tab measures how efficiently a harness drives a fixed model, not whether the model remembers a public bug. It does mean the resolve counts are contaminated on both sides, and the audit files in the bundle name every URL.
Cost uses DeepSeek's published standing rate for V4 Pro, the model both agents ran. Both runs actually executed in DeepSeek's off-peak window, where every rate is exactly half, so real spend was half of every dollar figure here. The standing rate is used because it is the list price and it applies to both sides identically.
DeepSeek charges nothing for a cache write. Superbrain's usual structural advantage over OpenCode — 39 percent fewer cache writes on the Terra tab, where a write is priced 12.5 times a read — is therefore worth exactly zero here. This tab is the harder comparison for us, and it is the one where the harness work of the last week shows up instead.
Superbrain takes about 8 percent longer in wall clock: 13m13s against 12m15s across the ten.
Superbrain issues more tool calls to get there. It costs less anyway, because a smaller share of those calls are wasted and each one carries more of the page it asked for.
An earlier build of Superbrain lost this comparison on the same ten bugs: 8 of 10 at 2.5 times OpenCode's cost per resolved instance. That run and this one are both in the repository. The difference is harness work, not a different model or a different set.
How it ran
Everything held identical except the agent.
The GitHub issue text, word for word, plus one line saying do not edit the tests. Same string to both agents.
DeepSeek V4 Pro, both sides on both sides, pinned, both billed to the same API key.
A fresh django checkout per bug, at the commit before the fix landed, with no later history in it.
The fix for every one of these bugs is public, and both agents were free to go and find it — the same freedom, on both sides. What is being measured is how efficiently each harness drives the model, not whether the model remembers the bug. Every transcript was audited and the bundle names each URL either agent reached for.
The official SWE-bench grader, in Docker. A bug counts as fixed only if the failing tests pass and the passing ones still do.
One rate card for both. We do not use either agent's self reported cost, because they price cached tokens differently.
Every bug, in numbers
The charts above, printed: outcome, billed tokens and cost per issue for both agents. Same data, nothing rounded away.
| Issue | Superbrain | Tokens | Cost | OpenCode | Tokens | Cost |
|---|---|---|---|---|---|---|
| django-10914Uploaded files got inconsistent permissions | Fixed | 99k | $0.0229 | Fixed | 224k | $0.0650 |
| django-10924A form field rejected a computed folder path | Fixed | 379k | $0.0637 | Fixed | 433k | $0.0833 |
| django-11001Multi line SQL broke query ordering | Fixed | 210k | $0.0494 | Fixed | 205k | $0.0453 |
| django-11019Combining form widgets warned about the wrong files | Fixed | 539k | $0.1187 | Fixed | 836k | $0.1488 |
| django-11039Migration output wrapped in a transaction it should not have | Fixed | 120k | $0.0206 | Fixed | 51k | $0.0133 |
| django-11049A duration field printed the wrong format in its error | Fixed | 132k | $0.0387 | Fixed | 191k | $0.0406 |
| django-11099A username validator accepted a trailing newline | Fixed | 119k | $0.0209 | Fixed | 58k | $0.0131 |
| django-11133HTTP responses mangled memoryview content | Fixed | 188k | $0.0542 | Fixed | 137k | $0.0287 |
| django-11179Deleting an object left a stale primary key | Fixed | 172k | $0.0400 | Fixed | 51k | $0.0136 |
| django-11283Permission migration crashed on existing rows | Fixed | 279k | $0.0788 | Fixed | 374k | $0.0720 |
Check it yourself
We ran this, so do not take our word for it. Rerun the official grader against our patches and see if you get the same 10 and 10.
python -m swebench.harness.run_evaluation --dataset_name princeton-nlp/SWE-bench_Lite --split test --predictions_path predictions/superbrain.jsonl --run_id verify
76 KB. Patches, grader output, container logs, per bug usage and the isolation audit. No agent transcripts, since those carry both companies' internals.