Nexrall Code — SWE-bench Verified
Pending — first run not yet published
We evaluate Nexrall Code (the nex CLI, powered by the Nexrall agent loop) against SWE-bench Verified, a set of 500 real-world GitHub issues with maintainer-verified fixes and hidden test suites. Every run below uses the official, unmodified SWE-bench evaluation harness — the same Docker-based grading pipeline used to produce the public SWE-bench leaderboard — so a % resolved shown here means the same thing it would mean on that leaderboard.
Why isn't this on swebench.com? As of November 2025, the official SWE-bench Verified leaderboard only accepts submissions from academic/ research-institution authors with a peer-reviewed or arXiv publication. Nexrall Code is a commercial product (Maxrall, Inc.), so we publish our own results here instead — using the exact same open-source harness, with full predictions and evaluation logs available for independent verification. See Methodology for exactly what that means in practice.
Latest result
| Model | Resolved | % Resolved | Instances | Run date | Report |
|---|---|---|---|---|---|
claude-sonnet-5 | — | — % | 500 | — | Run pending |
This table is updated after each full run completes. See Reproducing these results for the raw predictions, per-instance logs, and the exact harness invocation used to produce it.
Prior runs
None published yet — this is the first evaluation cycle.
Technical report
A short technical report describing Nexrall Code's agent architecture (tool loop, permission model, context management) as evaluated in this run will be linked here once the full 500-instance run completes.
— Henry Nguyen, Maxrall, Inc.