Skip to content
Blog

Beyond Pass Rate: Harness Evals, and Their Blind Spots

Anant Goel

Whatever lab you get tokens from likely gives you an agent harness along with it. The labs building these harnesses have a level of model access that nobody else has, and therefore often the model is trained inside the harness.

Switching harnesses always involves trade-offs. When you bring tokens into Delta from other tools, what are you giving up?

The current thing™ in agent evaluation is to rely on a subset of popular evals to compare the pass rate and cost of models inside their native harnesses. On the models we tested, Delta matches or exceeds the native harness on accuracy for all of them, while costing 0.80×–1.12× as much per pass depending on the model. Delta was cheaper per pass on Sol, roughly on par on Fable, and slightly more expensive on Astra and Opus. Given we use one system prompt across all models, I expected some model families to perform better than others, but instead I found both accuracy and costs for Delta to be competitive with their native harnesses with no directional result.

Models do as well in Delta as in their own harness

What changes when each model runs in Delta instead of its native harness.

Tasks solved
← native betterDelta better →
GPT-5.6 Solvs Codex
GPT-6 Astravs Codex
Claude Opus 5vs Claude Code
Claude Fable 5.12vs Claude Code
Cost per pass
← native betterDelta better →
GPT-5.6 Solvs Codex
GPT-6 Astravs Codex
Claude Opus 5vs Claude Code
Claude Fable 5.12vs Claude Code

Swipe horizontally to explore the full charts.

Source: Terminal-Bench 2.1, 29 tasks (25 for Fable) with the same reasoning effort in both harnesses.

A similar trend continues with open-weight models such as Kimi K3 [6], where Delta exceeds popular harnesses on accuracy while being competitive on cost.

Delta solves the most tasks at lower cost

Each dot is one harness running Kimi K3 on the same 30 tasks.

Swipe horizontally to explore the full charts.

Sources: Delta from our own Harbor runs of the 30 FrontierHarness tasks (Sept 23, 2026). All other harnesses from results published at frontierharness.org (Aug 22, 2026).

Time to pop the champagne, right? For us, not yet. These numbers don’t describe any of the reasons we at Zed love using Delta, or why we even built it. Delta is collaborative; it keeps the conversation and the code in one document and keeps you in the loop while the software is written.

Because public evals can’t showcase how our harness feels different, I want to explore more of what evals tell us today, and what they can’t yet capture.

Pass rates can’t reflect how code looks

Pass rates do matter, but the metric has known limitations. We could tell you Delta has a better pass rate than Codex for GPT-5.6 Sol, but this does not convey much of the human experience of using either tool. Pass rate can’t measure what the code looks like, how many files the agent touched, or whether a teammate could review it.

At Zed, each eval rollout is a thread we can open and inspect in Delta itself. This makes it easier to dig into the behavior of any model in Delta.

Let’s take a look at an example together and compare it against the same model completing the same task in its native harness:

In code-from-image from Terminal-Bench [7], the agent gets a picture of a pseudocode snippet, has to implement it and write the value it prints to output.txt. The snippet hashes the image file, then hashes that digest together with its first 10 bytes and a salt.

GPT-5.6 Sol in Codex worked out the answer with a one-off command in its shell:

node -e "const fs=require('fs'),c=require('crypto'); const b=fs.readFileSync('/app/code.png'); const h0=c.createHash('sha256').update(b).digest(); const H=c.createHash('sha256').update(Buffer.concat([h0,h0.subarray(0,10),Buffer.from('0000TBENCH-SALT')])).digest('hex'); process.stdout.write(H+'\n')"

Then it wrote the 64-character result into output.txt, and that file is all it left behind. The implementation the task asked for only ever existed inside that one command.

The same model in Delta wrote the same output.txt, and next to it solution.py:

from hashlib import sha256
from pathlib import Path


salt = b"0000TBENCH-SALT"
image_bytes = Path(__file__).with_name("code.png").read_bytes()
h0 = sha256(image_bytes).digest()
result = sha256(h0 + h0[:10] + salt).hexdigest()

print(result)

Before finishing, it ran python3 solution.py and compared the output against output.txt. A teammate opening this change can read the logic, rerun it, and see where the number came from.

We see a similar difference when using Claude Fable 5.1 in Claude Code and Delta. Neither Claude Code nor Delta left a script this time. In both, Fable worked out the hash in a throwaway Python snippet and checked it against a hint the task gives. The difference is the last step. Claude Code wrote the file by pasting the answer back in:

printf 'bee26a133f103b9ecda444c70ec22cafef6e31a3de7af6d047974dc90ce3defe\n' > /app/output.txt

In Delta, the code that hashed the image also wrote the file:

import hashlib
h0 = hashlib.sha256(open('/app/code.png','rb').read()).digest()
H = hashlib.sha256(h0 + h0[:10] + b'0000TBENCH-SALT').hexdigest()
open('/app/output.txt','w').write(H + '\n')
print(H)

If you’re just looking at the eval metrics, you wouldn’t notice this difference, but if you dig into the code, they’re different styles of writing code. In Delta, we don’t fold away the agent’s output from you the way other harnesses do, so we care deeply about how that output appears. Delta’s collaborative nature extends not just between two humans, but also to the human-agent interactions.

(Tool calls) failing upwards

Another example of something invisible to the evals yet obvious (and painful) to a developer is tool call failures. If an agent takes 5 attempts to make a tool call instead of just 1, are both equally correct? The eval scores would say yes, but users can feel the pain of the trade-offs made in the harnesses [1].

Claude Code, for example, has high reported failure rates. When I use Claude Code, I don’t see any of these failures because the harness silently absorbs them, but I feel the time I spend waiting on repeated requests. Delta and Claude Code are both getting the same tokens at the same speed, and Delta’s failure rate is ~1.5% on Anthropic’s models. The difference is that we don’t have to hide failures to get there, because we have DeltaDB.

Instead of hiding and retrying failed tool calls, we use our increased context to do something different (and better, we believe). Because Delta runs on DeltaDB, our harness knows exactly which version of a file the model read and what a past agent did, and we can use that intent and history to make better inferences that avoid failures in the first place. It’s not that we won’t ever fail, but richer context lets our harness be wiser without folding failures away.

Indulge me in an example: an edit that matches two places in the same file. Between the model reading the file and issuing a patch for it, the file can change: a formatter runs, the model’s own earlier edit shifts the lines, or a teammate types in the same buffer. Here’s how Delta uses its increased context to apply that tool call successfully:

How Delta handles ambiguous edits

When a patch matches two lines, Delta knows which one the model meant.

No retry, no re-read, no prompt guidance spent explaining any of it. And if the teammate had changed that line itself, Delta wouldn’t guess: it tells the model the content is stale and to read the file again.

The art of arranging tool calls

Even on successful tool calls, the ordering of those tool calls make a difference in the perceived speed of a model. Say you ask an agent to build a new feature (a pretty common way we use Delta ourselves). It will usually start by making a handful of tool calls to gather context around your codebase. You could arrange these tool calls in parallel, sequentially or in some combination of the two. Those choices cost you both time and tokens, but evals don’t score any of that. More tool calls per turn means fewer turns to finish the same task, and for you that means it feels a little faster and costs a little less.

Recently, Codex introduced async tool calls as a way to get the best of both worlds. However, the harness still plays a role in steering the model towards efficient behaviors.

Let’s look at the count-dataset-tokens task with GPT-6 Astra in Codex and Delta. The agent has to count the tokens in the science domain of a Hugging Face dataset, using a specific tokenizer, and write the number to answer.txt. Both harnesses pass this task on the eval.

  • Codex, 14 requests. It opens with a file search and a web search at once, and later installs packages alongside another command. In five of its requests it was checking on commands that were still running; Delta’s terminal waits for a command to finish instead.
  • Delta, 7 requests. It reads the dataset’s README while checking which Python packages are installed. In the next request it installs the missing packages while fetching the dataset’s file listing and the tokenizer’s config. Then it counts, writes the answer, and checks it.

Same task, half the requests

GPT-6 Astra on Terminal-Bench's count-dataset-tokens. Both runs passed.

Codex
110 seconds · 14 requests · 7 with several tool calls · 5 checking on commands
Delta
78 seconds · 7 requests · 2 with several tool calls
tool call
waiting on model
checking on a command
final answer

Swipe horizontally to explore the full charts.

We’ve found that this difference meaningfully affects wall time and cost, and that it varies across harnesses over many tasks, not just this one. Delta is built for people collaborating in real time, and people shouldn’t have to wait on machines when they don’t need to. Every request they sit through is time spent waiting, which is why the way a harness steers tool calls matters so much to us.

What’s next

So far, our focus for Delta’s agent harness has been to ensure we’re just as accurate and efficient as the harnesses we’re asking users to give up to work in Delta. That’s not enough; we want to design evals that communicate the reasons we reach for Delta (even for side projects where we pay for our tokens). Evals have always been a good way to imbue both the model and the harness with one’s judgment. In the next few months, we’ll standardize and publish deeper eval metrics (no more vanity metrics!), on a regular release cadence, to reflect the continuous and iterative process of building a harness. Our next update to this set of chosen tasks will shift focus to collaborative tasks that represent how users get work done in Delta.

Appendix

Why these evals?

Over the course of building Delta, I’ve spent a lot of my time using evals to understand how models behave in it compared to other harnesses. This question has also recently become part of AI research and discourse more broadly [2, 3, 4], with HarnessTax [5] and FrontierHarness [6] leading that conversation. We decided to use the same evals from these projects, drawn from Terminal-Bench [7] and DeepSWE [8], to evaluate Delta’s agent harness, to ensure there were public points of comparison and our results were not being artificially inflated by internal evals.

References

  1. Armin Ronacher. Better Models: Worse Tools. July 2026.
  2. Naman Vats and Oleg Golev. The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation. arXiv:2607.22585, 2026.
  3. Mohsen Arjmandi. Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite. arXiv:2609.11987, 2026.
  4. Sydney Lewis. Same Model, Different Harness: Different Coding-Agent Results. arXiv:2608.26218, 2026.
  5. Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica, and Matei Zaharia. HarnessTax: How Much Does Harness Matter for Coding Agents? 2026.
  6. Shilin Zhu and Shiqi Mei. Introducing FrontierHarness Eval. Runta, September 2026. Results at frontierharness.org.
  7. Mike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, et al. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces. The Fourteenth International Conference on Learning Representations (ICLR), 2026.
  8. Wenqi Huang, Charley Lee, Leonard Tng, and Serena Ge. DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks. 2026.

Footnotes

  1. Total cost of all tasks, including failed ones, divided by the number of tasks solved. FrontierHarness reports this figure as “median cost per task”, but it is not a statistical median, so we’ve renamed it to describe what it measures. ↩

  2. Fable has fewer tasks because its security classifier is triggered on some of the tasks. ↩