Methodology
What we run
We use the harness in nexrall-code/packages/bench/swebench (the same repository that ships the nex CLI), which has two phases:
- Inference (
run_inference.py) — for each of the 500 SWE-bench Verified task instances: clone the target repository at the task's pinnedbase_commit, runnex -p "<issue text>" --output-format jsonin headless mode against it, and capturegit diffas the candidate patch. - Evaluation (
evaluate.sh) — hand the resultingpreds.jsonlto the official, unmodifiedswebench.harness.run_evaluationmodule from SWE-bench/SWE-bench. This is the same grading code the public leaderboard uses: it builds a per-repo Docker container, applies the candidate patch, runs the task's hiddentest_patch, and checks the specificFAIL_TO_PASS/PASS_TO_PASStest IDs. We do not modify this code in any way.
What the agent sees — and doesn't
- The agent receives only
problem_statement(the public GitHub issue text). It never receivestest_patch(the hidden tests used to grade it). - It also never receives
hints_text— maintainer/PR discussion that often gives away the actual fix. The official SWE-bench submission checklist requires "Does not use thehintsfield"; we hold ourselves to the same rule even though we aren't submitting to that leaderboard. - The agent has a
fetch_urltool it can use to browse the web. For every benchmark run we setNEXRALL_FETCH_BLOCKLIST=github.com,...(and its API/raw-content/CDN subdomains), which makesfetch_urlrefuse any request to those hosts. This closes the obvious way an agent could otherwise look up the real PR that already fixed the issue it's being scored on — the same requirement the official checklist places on leaderboard submissions ("has taken steps to prevent lookup of SWE-bench solutions via web-browsing").
Attempt policy
Every result on this site is pass@1: one attempt per task instance, no retries, no best-of-k selection. run_inference.py runs each instance_id exactly once by construction — there is no code path that attempts an instance twice and reports the better outcome.
What we don't do
- We do not use SWE-bench test knowledge (
FAIL_TO_PASS/PASS_TO_PASS) to guide the agent — those are only used by the harness, after the fact, to grade the patch the agent already committed to. - We do not hand-pick which instances to run or exclude any instance from the published count. A run is either the full 500-instance Verified split or is explicitly labeled as a smaller "smoke test" subset.
- We do not alter the evaluation harness's pass/fail logic.
Cost & infrastructure
Each run executes on a single Linux x86_64 box (typically an AWS EC2 c6i.4xlarge). Inference cost (the actual model API spend across 500 multi-turn agentic sessions) is reported alongside each result where available.
Why we don't submit to swebench.com
As of November 18, 2025, the official SWE-bench Verified and Multilingual leaderboards only accept submissions from academic teams or research institutions with a peer-reviewed publication or arXiv preprint, with at least one author affiliated with an academic institution or established research lab. Maxrall, Inc. is a commercial company; we do not meet that criterion, and we aren't going to misrepresent an affiliation to satisfy it.
We publish here instead, using the identical open-source evaluation harness, so the numbers mean exactly what they'd mean if they were on that leaderboard — and so anyone can independently reproduce them. See Reproducing these results.
Terminal-Bench 4.0
Terminal-Bench results use a separate harness — nexrall-code/packages/bench/terminal-bench — built on the official Harbor runner rather than a hand-rolled clone-and-diff loop, because Terminal-Bench grades live container state (files written, processes started, servers listening) rather than a single patch.
What we run
A custom Harbor agent (nex_terminal_bench.NexAgent, a BaseInstalledAgent subclass) installs nex inside each task's own Docker container and runs it headlessly:
nex --output-format stream-json --yolo --no-banner --model <model>piped the task instruction over stdin. Harbor's own verifier — a second, isolated container per task — then runs the task's hidden test suite against whatever state the agent left behind and records a {"reward": 0.0 or 1.0} outcome. We do not modify Harbor's environment, verifier, or scoring logic in any way.
What the agent sees — and doesn't
- The agent receives only the task's public instruction text — never the verifier's hidden test suite.
- Every run sets
NEXRALL_FETCH_BLOCKLISTto covertbench.ai,www.tbench.ai, andhub.harborframework.com— the Terminal-Bench site and the Harbor Hub host that mirrors the dataset — so an agent can't look up the benchmark's own website/repo mid-task (the official leaderboard's own submission rules forbid this). The blocklist is deliberately narrower than a blanketgithub.com/huggingface.coban: tasks likemteb-leaderboardandcount-dataset-tokenspoint the agent at huggingface.co content as part of the task, and blocking it would fail the task through no fault of the agent. --yolo(auto-approve) is required because there is no TTY inside a Harbor task container to answer a permission prompt. The one guardrail we keep even under blanket auto-approval:NEXRALL_ALLOW_DESTRUCTIVEis never set, so the agent still cannot run a DB-drop/force-push/terraform destroy-class command — no Terminal-Bench 4.0 task's verifier requires one.
Terminal-Bench leaderboard status
Not yet leaderboard-eligible, tracked deliberately rather than silently:
- ATIF trajectories — done.
nex_terminal_bench/atif.py'sconvert_stream_json_to_trajectory()convertsnex'sstream-jsonevent log into a real ATIF-v1.7Trajectory, wired intonex_agent.py'spopulate_context_post_run(SUPPORTS_ATIF = True). This is no longer the gap it was under 2.1 — the converter ships and the adapter reports ATIF support. - A completed qualifying run — not yet done. Submission requires ≥5 trials per task (
k=5) across all 66 tasks, uploaded publicly to Harbor Hub, then run through the repo'slbpipeline (lb filter→lb metadata→lb open-prs) againstharbor-framework/terminal-bench. - Community submissions closed — the current blocker. As of this writing, Terminal-Bench 4.0's own
leaderboard/SUBMIT.mdstates "Community submissions are currently closed for Terminal-Bench 4.0. Only submissions run by the maintainers will be added to the leaderboard at this time." — the same posture 2.1 had at end-of-life.
Until a full run is done and maintainers reopen submissions, Terminal-Bench results published here are our own reproducible numbers — same posture SWE-bench Verified is in on this site: real numbers, real artifacts, just not (yet) a leaderboard entry.