Skip to content

Nexrall Code — SWE-bench Verified ​

Pending — first run not yet published

Looking for terminal-task results instead of a single-patch benchmark? See Terminal-Bench 4.0.

We evaluate Nexrall Code (the nex CLI, powered by the Nexrall agent loop) against SWE-bench Verified, a set of 500 real-world GitHub issues with maintainer-verified fixes and hidden test suites. Every run below uses the official, unmodified SWE-bench evaluation harness — the same Docker-based grading pipeline used to produce the public SWE-bench leaderboard — so a % resolved shown here means the same thing it would mean on that leaderboard.

Why isn't this on swebench.com? As of November 2025, the official SWE-bench Verified leaderboard only accepts submissions from academic/ research-institution authors with a peer-reviewed or arXiv publication. Nexrall Code is a commercial product (Maxrall, Inc.), so we publish our own results here instead — using the exact same open-source harness, with full predictions and evaluation logs available for independent verification. See Methodology for exactly what that means in practice.

Latest result ​

ModelResolved% ResolvedInstancesRun dateArtifacts
claude-sonnet-5-5—— %500—Run pending

This table is updated after each full run completes, by pasting the row publish_result.py prints once a run is published. See Reproducing these results for what "Artifacts" links to.

How a result gets published ​

Following the same split every public SWE-bench submission uses (SWE-bench/experiments, Multi-SWE-bench): small, diffable artifacts live in the nexrall-code git repo where anyone can browse them on GitHub; large per-instance execution logs are not embedded in this site — they're uploaded to a public object store and linked, exactly like Multi-SWE-bench hosts its trajectories on Hugging Face rather than building a log viewer.

Each row's Artifacts column links to three things:

  • predictions — preds.jsonl, the raw model-generated patch for every task instance, committed to nexrall-code under packages/bench/swebench/results/<model>/<run_id>/.
  • report — report.json, the official harness's own aggregate output (resolved / unresolved / error counts), same location.
  • logs — logs.tar.gz on benchmark-logs.nexrall.com: the complete logs/run_evaluation/<run_id>/ tree the official harness produced — patch.diff, report.json, and test_output.txt for every one of the 500 instances — plus a .sha256 checksum so you can verify the archive you downloaded matches what was published.

Prior runs ​

None published yet — this is the first evaluation cycle.

Technical report ​

A short technical report describing Nexrall Code's agent architecture (tool loop, permission model, context management) as evaluated in this run will be linked here once the full 500-instance run completes.

— Henry Nguyen, Maxrall, Inc.

Built by Maxrall, Inc. Not affiliated with or endorsed by the SWE-bench project.