jkchang.uk
Article · 12 Aug 2026Agentic systems · Context engineering8 min read

Own the harness

The model is the engine, and the harness around it does the work. Renting that harness from a vendor buys a fast demo and gives away the part that compounds, starting with memory.

When an agent fails, most teams reach first for a better model. The demo loops on a tool call, or loses track of the task halfway through, or gives a confident answer from the wrong document, and the conversation turns to whether the next release will be clever enough to stop. Occasionally it is. More often the failure has little to do with how intelligent the model is and a lot to do with how it is steered: what it is told, what it can see, which tools it has, what it remembers from the last step, and what is supposed to happen when something goes wrong.

I put this as a formula, because it keeps me honest about where the engineering is:

Agent = model + harness

The model is the engine. The harness is everything around it: the rules and instructions, the tools and their contracts, the workflow that decides what happens next, the memory that carries context between steps and sessions, and the checks that decide whether the work is done. Building agents for legal work showed me that the model is the replaceable part, so I keep models swappable on purpose and put the engineering into the harness.

The wrong-document failure is one I have dealt with myself. An agent answered confidently from the wrong document, or from an out-of-date version of the right one, and I don't think a stronger model would have helped, because it would have answered just as confidently from the same material. The fix was in the harness, in what that step retrieved and what it was allowed to see. Once that changed, the same model gave the right answer.

If the harness is where the value sits, then who owns it matters, and that is where I part company with much of the current agent market.

The case for renting is real

Every major model vendor now sells some form of harness as a service. Anthropic launched Claude Managed Agents in public beta in April: a hosted harness with sandboxing, tool execution, state and scoped permissions, where you hand over a task and get a result back. OpenAI's Agents API has the same shape, with the agent running in an OpenAI-hosted sandbox. The pitch is simple, and it is accurate. You get a working agent in ten minutes without building a loop, a sandbox, a state store or a permission model, and a team that has never built an agent can show one to its leadership by Friday.

For a prototype I would use one myself. The engineering inside those services is good, and building a sandbox or a state store from scratch just to find out whether an idea is worth pursuing is a poor use of anyone's time. The trouble starts when the demo is approved and the same foundation is asked to carry a business process, because the properties that made the demo fast are the ones that make production hard. I see five traps, and I'll take them in rising order of cost.

A rented harness and an owned harnessLeft: in a rented harness the loop, tools, memory and model all sit inside a vendor-hosted box, and the team's only lever is the prompt. Right: in an owned harness the workflow, tools, checks, memory and traces sit in the team's own code and stores, and the workflow calls a model that can be swapped for another. prompt calls reads, writes calls or swaps to Your team Loop Tools Memory and history Model Workflow Tools Memory and traces Checks Model A Model B Rented HARNESS AS A SERVICE Owned HARNESS YOU BUILD VENDOR-HOSTED one lever: the prompt YOUR CODE AND STORES every part readable, changeable, portable
Fig. 1 On the left, everything that decides behaviour sits inside the vendor's box, and the prompt is the only thing the team can change. On the right, the same parts sit in the team's own code and stores, and the model is one replaceable component.

You can't fix a loop you can't see

The first trap is the black box. A hosted harness decides how the loop runs: when the model is called again, how tool results are fed back, when the run stops and what happens on an error. Suppose an agent gets stuck, calling the same search tool ten times with slightly reworded queries, or retrying an action that will never succeed. The fix belongs in the loop. In a harness you own, you read the trace, find the step that went wrong and change what caused it: put a budget on tool calls, deduplicate repeated queries, change the stop condition, or narrow what that step can see. In a harness you rent, the loop isn't yours to change, so you adjust the prompt and hope.

That leads to the second trap, which is what debugging becomes when the prompt is your only lever. Every fix lands in the same place. You add a sentence to stop one failure, and that sentence is read on every run, including the hundreds it was never written for, so it quietly shifts behaviour that used to work. You fix one error and break three things that nobody notices until a user does. The system doesn't get smarter with each fix; it gets more fragile, because the corrections pile up in one shared block of text that nobody can reason about any more.

Owning the harness doesn't make failures go away, but it does let you put each fix where the failure is. A malformed tool call gets validation on that tool. A wrong answer from stale context gets a change to what that step retrieves. An unsafe action gets a check before it runs. Each fix stays local, and because the traces are yours, you can replay yesterday's runs against today's harness and see what changed before your users do.

A 60% demo and a 95% system are different architectures

The third trap is expecting the quick-start foundation to scale. What gets an agent working 60% of the time is a different architecture from what gets it to 95%, and no amount of tuning closes that gap.

Some arithmetic shows why. A workflow of ten steps, each succeeding 95% of the time, finishes cleanly only about 60% of the time, because 0.95 to the tenth power is 0.60, and real business processes have more than ten steps. Reliability at that length comes from structure, not from a better single loop. You break the work into steps with defined inputs and outputs, check each result before the next step builds on it, keep state so that a failed run can resume instead of starting again, route risky actions to a person for approval, and log enough to know afterwards what happened. None of these pieces is exotic, but they all live in the harness, and a quick-start framework is designed so that you don't have to think about them. I traced that progression, from a single loop to a governed graph, in From loops to graphs.

A team that builds a complex process on a hosted demo harness is taking on technical debt from the first day. Each workaround for a missing piece of structure is more glue around a box they can't open, and when they finally need the structure, they find they have to rebuild the system rather than extend it.

You move at the vendor's speed

The fourth trap is slower and easier to miss. When your core workflow is written against one vendor's framework, you can only do what that framework supports, and only once it supports it. If you need a checkpoint the service doesn't expose, a tool pattern it doesn't allow, or another provider's model for one step, you wait or you work around it. And because a vendor's hosted harness runs that vendor's models, the one part of the system I deliberately keep replaceable becomes the part you can't replace.

Elsewhere in software this would be a familiar lock-in trade, and often an acceptable one. What makes it expensive here is the pace. Models, tool protocols and agent patterns have been turning over in months rather than years, and a team that can try a new model on one step on a Tuesday afternoon learns faster than a team filing a feature request. Flexibility is what you give up, and right now it is worth a great deal.

Memory is the part you can't get back

The fifth trap is the most dangerous, and the one I care about most. An agent that works on real tasks accumulates things: the context of ongoing work, the history of what was asked and what was done, corrections from the people who reviewed it, and specialised knowledge of how this organisation does this kind of task. That accumulated memory is what makes the hundredth run better than the first. I think of the moat for AI in a business as workflow plus knowledge, and an agent's memory is where the two meet.

When the harness is rented, the memory lives where the harness lives. Even if a vendor lets you export it, the decisions that give memory its value (what is kept, how it is structured, what gets recalled into which step, what is forgotten) follow the vendor's design, tuned for their typical customer rather than for your domain. Your traces, which are also the raw material for evaluating and improving the agent, sit in someone else's format. Over time, knowledge that should have been compounding inside your organisation compounds inside a product you rent, and the cost of leaving grows with every run.

I don't think most teams make this trade knowingly. It arrives as a default, a convenient session store in the quick-start, and by the time anyone asks where the agent's knowledge actually lives, a year of it is somewhere else.

What owning it means

Owning the harness doesn't mean building everything. I still call models through an API, and I would happily use a vendor's sandbox as a component, the way I'd use any managed service: something my harness calls, behind an interface I control, that I could swap out if I had to. In The third generation of skills I argued that agents will increasingly write their own code for the task in front of them, and a hosted sandbox is a sensible place to run that code, because the code is disposable. The loop that decides what to run, and the memory it learns from, are not. So the line I draw is around the parts that decide behaviour and the parts that accumulate value. The loop, the workflow, the tool contracts, the checks, the memory and the traces should be yours, in your own code and stores, where you can read them, change them and carry them to the next model.

That costs more at the start. The ten-minute demo becomes weeks of engineering before the first useful run, and the first version may well look worse than the vendor's. For me the hardest part was state and memory: knowing where a run had got to, what it should carry forward, and how to resume it after a failure without starting again. It is no coincidence that this is also the part a hosted service makes easiest to hand over. What the extra work buys is a system that improves where you can see it, and knowledge that stays with the organisation that earned it.

So the next time an agent fails, look at the harness before blaming the model, and if you can't look at it, that is the first problem to solve. If AI is going to be a real competitive advantage, you have to build and own the harness. Don't outsource the steering wheel, and never give away your knowledge.

Jiakang Chang · Principal Software EngineerAll articles