Reproducing these results
Every result on this site links to three artifacts — the same split used by the official SWE-bench/experiments leaderboard repo and by Multi-SWE-bench: small, diffable files in git; large per-instance logs in a public object store, linked rather than embedded in this site.
- Predictions (
preds.jsonl) — the raw model-generated patch for every task instance, including any that errored out. Committed to thenexrall-coderepo atpackages/bench/swebench/results/<model>/<run_id>/preds.jsonl. - Result summary (
report.json) — the official harness's own aggregate output: resolved / unresolved / error counts. Same location as above. - Evaluation logs (
logs.tar.gz) — the completelogs/run_evaluation/<run_id>/tree exactly as produced by the official harness: per-instancepatch.diff,report.json(pass/fail againstFAIL_TO_PASS/PASS_TO_PASS), andtest_output.txt(the actual test run inside the grading container). Hosted publicly athttps://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz.
Verifying a logs archive hasn't been altered
Every logs.tar.gz is published alongside a logs.tar.gz.sha256 checksum file at the same URL. Before trusting the contents, verify it:
curl -fsSL https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz -o logs.tar.gz
curl -fsSL https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz.sha256 -o logs.tar.gz.sha256
shasum -a 256 -c logs.tar.gz.sha256 # or: sha256sum -c logs.tar.gz.sha256
tar xzf logs.tar.gzThe checksum is generated at publish time (publish_result.py), before the archive is uploaded, and is never regenerated afterward — if the two don't match, the archive was corrupted or altered in transit and should not be trusted for verification.
Running it yourself
The harness is the same one Nexrall Code uses internally — no special access required beyond a Nexrall account (for CLI auth) and standard SWE-bench infrastructure (Docker, ~120GB disk, a Linux x86_64 box).
# 1. One-time box setup (installs Docker, the official SWE-bench harness, and
# the nex CLI via npm).
bash setup.sh
source ~/swebench-venv/bin/activate
export NEXRALL_TOKEN=... # headless CLI auth — from your Nexrall account
# 2. Smoke test (20 instances) — cheap, proves the pipeline end-to-end.
python3 run_inference.py --out preds20.jsonl --model claude-sonnet-5-5 --limit 20 --resume
bash evaluate.sh preds20.jsonl smoke20 4
# 3. Full 500-instance run.
python3 run_inference.py --out preds_full.jsonl --model claude-sonnet-5-5 --resume
bash evaluate.sh preds_full.jsonl full-claude-sonnet-5-5 8Because Phase B (evaluate.sh) invokes the unmodifiedswebench.harness.run_evaluation module from SWE-bench/SWE-bench, re-running it on our published preds.jsonl will reproduce the same pass/fail outcome per instance (modulo the harness's own documented flakiness on a small number of known-flaky instances, same as it would for any submission on the official leaderboard).
Terminal-Bench
Terminal-Bench results use a separate harness at nexrall-code/packages/bench/terminal-bench, built on the official Harbor runner. Same split as above:
- Per-task rewards (
summary.json) — committed tonexrall-codeatpackages/bench/terminal-bench/results/<model>/<run_id>/summary.json. - Run index (
index.json) — same location, includes the logs URL and checksum. - Job artifacts (
logs.tar.gz) — the full Harbor job directory (every trial'sresult.json, agent output, verifier output), hosted athttps://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gzwith a.sha256checksum alongside it, verified the same way as the SWE-bench archive above.
Running it yourself
Requires a Nexrall account (for CLI auth), Docker, and Harbor (uv tool install harbor).
cd packages/bench/terminal-bench
bash setup.sh # installs uv, Harbor, this adapter package
source .venv/bin/activate
export NEXRALL_TOKEN=... # headless CLI auth — from your Nexrall account
# 1. Smoke test — 5 tasks, k=1 — proves the pipeline end-to-end.
bash run.sh --smoke
# 2. Full run — all 66 tasks, k=5 (leaderboard-eligible trial count).
CONCURRENCY=8 bash run.sh --fullBecause this invokes the unmodified Harbor runner against the unmodified terminal-bench/terminal-bench@4.0.0 dataset, re-running it will reproduce the same reward per trial (modulo Terminal-Bench's own documented environment flakiness on a small number of tasks with external network dependencies, same as it would for any Harbor-based submission).
Questions about a specific result
If a number here doesn't reproduce for you, or you'd like to run a random subset yourself to spot-check a published result, open an issue or reach out — contact details are on nexrall.com.