Skip to content

Nexrall Code — Terminal-Bench 4.0 ​

Pending — full run not yet published

We evaluate Nexrall Code (the nex CLI) against Terminal-Bench 4.0, a set of 66 realistic terminal-based engineering tasks — software engineering, security, data science, model training, and system administration — each graded by running a task-specific verifier against the live state of a sandboxed Docker container the agent worked in. Unlike SWE-bench (which grades a single git patch), Terminal-Bench grades the terminal session itself: an agent that never inspects a file it needed to read fails here even if it gets lucky with a patch.

Every run below uses the official, unmodified Harbor harness — the same runner that produces the public Terminal-Bench leaderboard — via a custom Harbor agent adapter that runs nex headlessly inside each task's container. See Methodology for exactly what that harness does and doesn't let the agent do.

Is this on the official Terminal-Bench leaderboard? Not yet. Two independent reasons, tracked deliberately rather than silently. First, the official leaderboard's automated validation requires every submitted trial to carry a real ATIF agent trajectory so maintainers can audit for reward hacking — our adapter produces one (SUPPORTS_ATIF = True), so this is done. Second, and the current blocker, community submissions are closed for Terminal-Bench 4.0 (the project states only maintainer-run submissions are added at this time), and we have not yet completed a qualifying k=5 run. Numbers on this page are our own reproducible results, not a leaderboard entry. See Methodology for the exact status.

Latest result ​

ModelTasks resolved% ResolvedTasksRun dateArtifacts
deepseek-v4-pro—— %66—Full run pending

This table is updated after each full (k=5, all 66 tasks) run completes, by pasting the row publish_result.py prints once a run is published. See Reproducing these results for what "Artifacts" links to.

Smoke test (in progress) ​

Before committing to a full 66-task × 5-trial run, we always validate the harness end-to-end on a handful of tasks first — cheap, fast, and catches infrastructure problems (missing system packages, auth, model connectivity) before they'd otherwise silently fail hundreds of trials. Validation runs use the same standard config as the planned full run — deepseek-v4-pro at max reasoning effort — and results will be summarized here once complete, before any full run is scheduled.

How a result gets published ​

Following the same small-artifacts-in-git / large-artifacts-in-object-store split the SWE-bench Verified page on this site uses:

  • Per-task rewards (summary.json) — the raw {"reward": 0.0 or 1.0} outcome for every trial, aggregated per task. Committed to the nexrall-code repo at packages/bench/terminal-bench/results/<model>/<run_id>/summary.json.
  • Run index (index.json) — resolved/total counts, pass@1, run date, and the logs URL/checksum, same location as above.
  • Job artifacts (logs.tar.gz) — the complete Harbor job directory for the run: per-trial result.json, agent output logs (agent/nex-output.jsonl), and verifier output for every one of the 66 tasks × 5 trials. Hosted publicly at https://benchmark-logs.nexrall.com/<model>/<run_id>/logs.tar.gz, with a .sha256 checksum alongside it — same verification flow as the SWE-bench artifacts, see Reproducing these results.

Prior runs ​

None published yet — this is the first evaluation cycle.

— Henry Nguyen, Maxrall, Inc.

Built by Maxrall, Inc. Not affiliated with or endorsed by the SWE-bench project.