Skip to content

Methodology ​

What we run ​

We use the harness in nexrall-code/packages/bench/swebench (the same repository that ships the nex CLI), which has two phases:

  1. Inference (run_inference.py) — for each of the 500 SWE-bench Verified task instances: clone the target repository at the task's pinned base_commit, run nex -p "<issue text>" --output-format json in headless mode against it, and capture git diff as the candidate patch.
  2. Evaluation (evaluate.sh) — hand the resulting preds.jsonl to the official, unmodified swebench.harness.run_evaluation module from SWE-bench/SWE-bench. This is the same grading code the public leaderboard uses: it builds a per-repo Docker container, applies the candidate patch, runs the task's hidden test_patch, and checks the specific FAIL_TO_PASS / PASS_TO_PASS test IDs. We do not modify this code in any way.

What the agent sees — and doesn't ​

  • The agent receives only problem_statement (the public GitHub issue text). It never receives test_patch (the hidden tests used to grade it).
  • It also never receives hints_text — maintainer/PR discussion that often gives away the actual fix. The official SWE-bench submission checklist requires "Does not use the hints field"; we hold ourselves to the same rule even though we aren't submitting to that leaderboard.
  • The agent has a fetch_url tool it can use to browse the web. For every benchmark run we set NEXRALL_FETCH_BLOCKLIST=github.com,... (and its API/raw-content/CDN subdomains), which makes fetch_url refuse any request to those hosts. This closes the obvious way an agent could otherwise look up the real PR that already fixed the issue it's being scored on — the same requirement the official checklist places on leaderboard submissions ("has taken steps to prevent lookup of SWE-bench solutions via web-browsing").

Attempt policy ​

Every result on this site is pass@1: one attempt per task instance, no retries, no best-of-k selection. run_inference.py runs each instance_id exactly once by construction — there is no code path that attempts an instance twice and reports the better outcome.

What we don't do ​

  • We do not use SWE-bench test knowledge (FAIL_TO_PASS/PASS_TO_PASS) to guide the agent — those are only used by the harness, after the fact, to grade the patch the agent already committed to.
  • We do not hand-pick which instances to run or exclude any instance from the published count. A run is either the full 500-instance Verified split or is explicitly labeled as a smaller "smoke test" subset.
  • We do not alter the evaluation harness's pass/fail logic.

Cost & infrastructure ​

Each run executes on a single Linux x86_64 box (typically an AWS EC2 c6i.4xlarge). Inference cost (the actual model API spend across 500 multi-turn agentic sessions) is reported alongside each result where available.

Why we don't submit to swebench.com ​

As of November 18, 2025, the official SWE-bench Verified and Multilingual leaderboards only accept submissions from academic teams or research institutions with a peer-reviewed publication or arXiv preprint, with at least one author affiliated with an academic institution or established research lab. Maxrall, Inc. is a commercial company; we do not meet that criterion, and we aren't going to misrepresent an affiliation to satisfy it.

We publish here instead, using the identical open-source evaluation harness, so the numbers mean exactly what they'd mean if they were on that leaderboard — and so anyone can independently reproduce them. See Reproducing these results.

Terminal-Bench 4.0 ​

Terminal-Bench results use a separate harness — nexrall-code/packages/bench/terminal-bench — built on the official Harbor runner rather than a hand-rolled clone-and-diff loop, because Terminal-Bench grades live container state (files written, processes started, servers listening) rather than a single patch.

What we run ​

A custom Harbor agent (nex_terminal_bench.NexAgent, a BaseInstalledAgent subclass) installs nex inside each task's own Docker container and runs it headlessly:

nex --output-format stream-json --yolo --no-banner --model <model>

piped the task instruction over stdin. Harbor's own verifier — a second, isolated container per task — then runs the task's hidden test suite against whatever state the agent left behind and records a {"reward": 0.0 or 1.0} outcome. We do not modify Harbor's environment, verifier, or scoring logic in any way.

What the agent sees — and doesn't ​

  • The agent receives only the task's public instruction text — never the verifier's hidden test suite.
  • Every run sets NEXRALL_FETCH_BLOCKLIST to cover tbench.ai, www.tbench.ai, and hub.harborframework.com — the Terminal-Bench site and the Harbor Hub host that mirrors the dataset — so an agent can't look up the benchmark's own website/repo mid-task (the official leaderboard's own submission rules forbid this). The blocklist is deliberately narrower than a blanket github.com/huggingface.co ban: tasks like mteb-leaderboard and count-dataset-tokens point the agent at huggingface.co content as part of the task, and blocking it would fail the task through no fault of the agent.
  • --yolo (auto-approve) is required because there is no TTY inside a Harbor task container to answer a permission prompt. The one guardrail we keep even under blanket auto-approval: NEXRALL_ALLOW_DESTRUCTIVE is never set, so the agent still cannot run a DB-drop/force-push/ terraform destroy-class command — no Terminal-Bench 4.0 task's verifier requires one.

Terminal-Bench leaderboard status ​

Not yet leaderboard-eligible, tracked deliberately rather than silently:

  1. ATIF trajectories — done. nex_terminal_bench/atif.py's convert_stream_json_to_trajectory() converts nex's stream-json event log into a real ATIF-v1.7 Trajectory, wired into nex_agent.py's populate_context_post_run (SUPPORTS_ATIF = True). This is no longer the gap it was under 2.1 — the converter ships and the adapter reports ATIF support.
  2. A completed qualifying run — not yet done. Submission requires ≥5 trials per task (k=5) across all 66 tasks, uploaded publicly to Harbor Hub, then run through the repo's lb pipeline (lb filter → lb metadata → lb open-prs) against harbor-framework/terminal-bench.
  3. Community submissions closed — the current blocker. As of this writing, Terminal-Bench 4.0's own leaderboard/SUBMIT.md states "Community submissions are currently closed for Terminal-Bench 4.0. Only submissions run by the maintainers will be added to the leaderboard at this time." — the same posture 2.1 had at end-of-life.

Until a full run is done and maintainers reopen submissions, Terminal-Bench results published here are our own reproducible numbers — same posture SWE-bench Verified is in on this site: real numbers, real artifacts, just not (yet) a leaderboard entry.

Built by Maxrall, Inc. Not affiliated with or endorsed by the SWE-bench project.