Skip to content

Reproducing these results ​

Every result on this site links to three artifacts — the same split used by the official SWE-bench/experiments leaderboard repo and by Multi-SWE-bench: small, diffable files in git; large per-instance logs in a public object store, linked rather than embedded in this site.

  • Predictions (preds.jsonl) — the raw model-generated patch for every task instance, including any that errored out. Committed to the nexrall-code repo at packages/bench/swebench/results/<model>/<run_id>/preds.jsonl.
  • Result summary (report.json) — the official harness's own aggregate output: resolved / unresolved / error counts. Same location as above.
  • Evaluation logs (logs.tar.gz) — the complete logs/run_evaluation/<run_id>/ tree exactly as produced by the official harness: per-instance patch.diff, report.json (pass/fail against FAIL_TO_PASS/PASS_TO_PASS), and test_output.txt (the actual test run inside the grading container). Hosted publicly at https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz.

Verifying a logs archive hasn't been altered ​

Every logs.tar.gz is published alongside a logs.tar.gz.sha256 checksum file at the same URL. Before trusting the contents, verify it:

bash
curl -fsSL https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz -o logs.tar.gz
curl -fsSL https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz.sha256 -o logs.tar.gz.sha256
shasum -a 256 -c logs.tar.gz.sha256   # or: sha256sum -c logs.tar.gz.sha256
tar xzf logs.tar.gz

The checksum is generated at publish time (publish_result.py), before the archive is uploaded, and is never regenerated afterward — if the two don't match, the archive was corrupted or altered in transit and should not be trusted for verification.

Running it yourself ​

The harness is the same one Nexrall Code uses internally — no special access required beyond a Nexrall account (for CLI auth) and standard SWE-bench infrastructure (Docker, ~120GB disk, a Linux x86_64 box).

bash
# 1. One-time box setup (installs Docker, the official SWE-bench harness, and
#    the nex CLI via npm).
bash setup.sh
source ~/swebench-venv/bin/activate
export NEXRALL_TOKEN=...          # headless CLI auth — from your Nexrall account

# 2. Smoke test (20 instances) — cheap, proves the pipeline end-to-end.
python3 run_inference.py --out preds20.jsonl --model claude-sonnet-5-5 --limit 20 --resume
bash evaluate.sh preds20.jsonl smoke20 4

# 3. Full 500-instance run.
python3 run_inference.py --out preds_full.jsonl --model claude-sonnet-5-5 --resume
bash evaluate.sh preds_full.jsonl full-claude-sonnet-5-5 8

Because Phase B (evaluate.sh) invokes the unmodifiedswebench.harness.run_evaluation module from SWE-bench/SWE-bench, re-running it on our published preds.jsonl will reproduce the same pass/fail outcome per instance (modulo the harness's own documented flakiness on a small number of known-flaky instances, same as it would for any submission on the official leaderboard).

Terminal-Bench ​

Terminal-Bench results use a separate harness at nexrall-code/packages/bench/terminal-bench, built on the official Harbor runner. Same split as above:

  • Per-task rewards (summary.json) — committed to nexrall-code at packages/bench/terminal-bench/results/<model>/<run_id>/summary.json.
  • Run index (index.json) — same location, includes the logs URL and checksum.
  • Job artifacts (logs.tar.gz) — the full Harbor job directory (every trial's result.json, agent output, verifier output), hosted at https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz with a .sha256 checksum alongside it, verified the same way as the SWE-bench archive above.

Running it yourself ​

Requires a Nexrall account (for CLI auth), Docker, and Harbor (uv tool install harbor).

bash
cd packages/bench/terminal-bench
bash setup.sh                         # installs uv, Harbor, this adapter package
source .venv/bin/activate
export NEXRALL_TOKEN=...              # headless CLI auth — from your Nexrall account

# 1. Smoke test — 5 tasks, k=1 — proves the pipeline end-to-end.
bash run.sh --smoke

# 2. Full run — all 66 tasks, k=5 (leaderboard-eligible trial count).
CONCURRENCY=8 bash run.sh --full

Because this invokes the unmodified Harbor runner against the unmodified terminal-bench/terminal-bench@4.0.0 dataset, re-running it will reproduce the same reward per trial (modulo Terminal-Bench's own documented environment flakiness on a small number of tasks with external network dependencies, same as it would for any Harbor-based submission).

Questions about a specific result ​

If a number here doesn't reproduce for you, or you'd like to run a random subset yourself to spot-check a published result, open an issue or reach out — contact details are on nexrall.com.

Built by Maxrall, Inc. Not affiliated with or endorsed by the SWE-bench project.