9 min left
0% read
A well-instructed agent with great context still deleted a developer's home directory. Here's why that failure sits at the harness layer, not the prompt, and what that actually means to build.
On August 26, 2026, a developer asked Claude to build a sandbox mechanism meant to stop AI agents from filling up his disk with temporary files. Claude wrote a script to detect whether other agents were still active before deleting their temp data. The review process flagged the approach twice and downgraded the model partway through. Then, while testing whether the sandbox actually worked, the agent ran a recursive delete against the developer's real home directory instead of the isolated test path. Roughly 700 gigabytes gone.
This wasn't an isolated event either. Anthropic's own Claude Code repository has at least five separately filed incidents describing the same failure class over six months: home directories or root paths getting wiped by an agent-constructed command that nothing in the surrounding system caught before it ran.
The instructions Claude was given were reasonable. The reasoning that led to the test wasn't obviously reckless. What failed was the layer around the model, the part that should have made a destructive command targeting a real home directory architecturally impossible to run. That layer has a name now: the harness. And the industry spent a lot of 2026 figuring out it's a genuinely different discipline from the two that came before it.
Prompt engineering is the oldest and narrowest of the three. It's the wording of a single request, the part a person writes by hand to get a better answer out of one model call. If a model misreads a clearly written instruction, that's a prompt problem, and it's usually visible immediately because the answer comes back obviously wrong.
Context engineering is broader. It's the discipline of deciding what information actually fills the model's context window before it answers: conversation history, retrieved documents, tool definitions, memory, application state. Good context engineering doesn't mean stuffing in everything available. More context that isn't relevant adds noise and makes the model more likely to miss what matters, not less. It means curating what the model sees for a given call.
Harness engineering is different from both, and it's the one that doesn't show up in the answer itself. It's the system that runs around the model across an entire task, not a single call: executing tool calls, deciding what's allowed to run without asking, catching a bad result before it reaches the next step, retrying failures, enforcing limits, and giving the whole thing a way to stop. A model with perfect instructions and perfect context can still fail here, because this layer governs what happens after the model decides what it wants to do, not what it's told or what it knows.
| Dimension | Prompt Engineering | Context Engineering | Harness Engineering |
|---|---|---|---|
| Controls | The wording of one request | What information the model sees | What the agent is allowed to do, and what catches it when it's wrong |
| Operates at | A single model call | A session or task | The full execution loop, across many calls |
| Failure looks like | The model misunderstands a clear instruction | The model confidently reasons over the wrong information | A well-instructed, well-informed agent still does something destructive or unrecoverable |
| Who notices first | Immediately, the output is visibly wrong | Often much later, the output looks plausible but is built on bad input | Sometimes never, until the one time a guardrail that didn't exist actually mattered |
If the rm -rf incident sounds like an edge case, a cluster of 2026 research on agent benchmarking shows the same gap shows up constantly, just measured in performance instead of catastrophe. Multiple papers have now documented that holding a model completely fixed and only changing the harness around it moves benchmark scores by a wide margin, often more than switching to a different model entirely.
On SWE-bench Pro, Claude Opus 4.5 scores 45.9% under a standardized evaluation scaffold called SEAL, and 55.4% under the Claude Code harness. That's roughly a 9.5 point swing with the exact same model weights. The Holistic Agent Leaderboard has reported single-model swings of up to nearly 48 percentage points on SWE-bench Verified Mini depending purely on scaffold choice. One analysis found a basic scaffold scoring 23% on SWE-bench Pro while an optimized version of the same setup scored over 45%, a 22-point gain from harness changes alone, no model upgrade involved. In one case, adding a single search subagent to an otherwise identical setup was enough to flip the ranking order between two different frontier models on the same benchmark.
The researchers behind one of these papers put it plainly: a benchmark score is the joint outcome of the model and the harness, but published leaderboards almost always report it as if it were a model property alone.
💡 Harness-driven swings routinely dwarf the 2 to 4 percentage point differences that get reported as meaningful progress between model generations. Reading "Model X scores 65% on SWE-bench" without knowing what harness ran it is, by this research, closer to meaningless than people tend to assume.
It's worth being specific about which layer failed, because it wasn't a prompt problem or a context problem. The model wasn't confused about what the user wanted, and it wasn't missing relevant information about the task. What was missing was a boundary the harness should have enforced: no architectural check that stopped a recursive delete from resolving to a real home directory, and no sandbox isolation strong enough to make the test environment's filesystem genuinely separate from the host's.
Related incidents reported in the same repository describe the same gap from different angles: a destructive command hidden inside a self-written helper script, a shell variable expansion that silently turned a scoped delete into a root-level one, a cleanup step that restored the real $HOME value right before the delete ran against it.
None of that is something a better prompt fixes. A better prompt doesn't stop a correctly-written, syntactically valid command from executing against the wrong path. That's what a harness is for.
Strip the diagram down and a harness is a small number of concrete responsibilities, not one big thing:
A practical version of the specific gap in the rm -rf case looks like this, a guard that runs before a destructive command executes, not after:
import os
import re
PROTECTED_PATTERNS = [
r"^rm\s+-rf\s+/\*?\s*$", # rm -rf / or rm -rf /*
r"^rm\s+-rf\s+\$?HOME\s*$", # rm -rf $HOME or rm -rf HOME
r"^rm\s+-rf\s+~/?\s*$", # rm -rf ~ or rm -rf ~/
]
def check_destructive_command(command: str, resolved_path: str) -> bool:
"""Returns True if the command is safe to execute."""
for pattern in PROTECTED_PATTERNS:
if re.match(pattern, command.strip()):
return False
if resolved_path in (os.path.expanduser("~"), "/"):
return False
return True
That's not a sophisticated piece of engineering. It's a few lines that check what a command actually resolves to before letting it run, instead of trusting that a reasonable-sounding instruction produces a safe result. The GitHub issues on this exact incident class recommend close to this: evaluate destructive commands after shell expansion, not before, and sandbox agent execution away from host mounts by default so that even a command that does go wrong lands somewhere recoverable.
Worth being honest that this isn't fully settled as a field yet. Most sources describe these three as nested layers, harness wraps context, context wraps prompt, where a strong harness still needs good context and a well-written prompt underneath it, it just adds leverage on top.
But the disagreement around that framing is real, not just semantic. Some writers treat harness engineering as a superset covering the other two rather than a separate layer stacked on top of them. And there's an emerging fourth term, sometimes called loop engineering or graph engineering, for the coordination layer when multiple agents or sub-agents work together, which some treat as part of the harness and others treat as its own discipline entirely. If you see slightly different diagrams in different places, that's why.
A good harness doesn't make the model smarter. It makes the system around the model survive the model being wrong, which it eventually will be, no matter how good the prompt or how well-curated the context.
The rm -rf incidents are a clean example precisely because nothing about them was a strange edge case. An agent testing whether a cleanup mechanism works is an extremely ordinary task. The failure wasn't exotic. The missing guardrail was.
If you're building anything that gives a model real tool access, the question worth asking isn't "is the prompt good enough" or "does it have the right context." It's: what happens the one time this agent does something destructive, and is there anything in the system, not the instructions, that catches it before it runs.
Harness engineering is the discipline of building the system that runs around a model across an entire task: executing tool calls, deciding what's allowed to run without asking, checking results before they reach the next step, retrying failures, enforcing limits, and giving the loop a way to stop. It governs what happens after the model decides what it wants to do.
Prompt engineering controls the wording of one request and operates on a single model call. Context engineering controls what information the model sees and operates across a session. Harness engineering controls what the agent is allowed to do and operates across the full execution loop — so it's the only one of the three that can stop a destructive action.
Holding the model fixed and changing only the harness moves scores by a wide margin. Claude Opus 4.5 scores 45.9% on SWE-bench Pro under the SEAL scaffold and 55.4% under the Claude Code harness — a 9.5 point swing on identical weights. The Holistic Agent Leaderboard has reported single-model swings of up to nearly 48 percentage points on SWE-bench Verified Mini from scaffold choice alone.
A harness boundary, not a better prompt. There was no architectural check stopping a recursive delete from resolving to a real home directory, and no sandbox isolation strong enough to keep the test filesystem genuinely separate from the host's. The command was syntactically valid and executed exactly as written.
Four concrete responsibilities: permissions (what runs without asking, and what never runs), tool execution plus a separate step that verifies results instead of trusting them, state that survives across the task rather than one call, and a bounded loop with a real exit condition instead of trusting the model to know when to stop.
Similar Topics