Skip to content

Nexrall Code — SWE-bench Verified

Pending — first run not yet published

We evaluate Nexrall Code (the nex CLI, powered by the Nexrall agent loop) against SWE-bench Verified, a set of 500 real-world GitHub issues with maintainer-verified fixes and hidden test suites. Every run below uses the official, unmodified SWE-bench evaluation harness — the same Docker-based grading pipeline used to produce the public SWE-bench leaderboard — so a % resolved shown here means the same thing it would mean on that leaderboard.

Why isn't this on swebench.com? As of November 2025, the official SWE-bench Verified leaderboard only accepts submissions from academic/ research-institution authors with a peer-reviewed or arXiv publication. Nexrall Code is a commercial product (Maxrall, Inc.), so we publish our own results here instead — using the exact same open-source harness, with full predictions and evaluation logs available for independent verification. See Methodology for exactly what that means in practice.

Latest result

ModelResolved% ResolvedInstancesRun dateReport
claude-sonnet-5— %500Run pending

This table is updated after each full run completes. See Reproducing these results for the raw predictions, per-instance logs, and the exact harness invocation used to produce it.

Prior runs

None published yet — this is the first evaluation cycle.

Technical report

A short technical report describing Nexrall Code's agent architecture (tool loop, permission model, context management) as evaluated in this run will be linked here once the full 500-instance run completes.

— Henry Nguyen, Maxrall, Inc.

Built by Maxrall, Inc. Not affiliated with or endorsed by the SWE-bench project.