jkchang.uk
Featured · 21 Jul 2026Agentic systems · Knowledge graphs8 min read

From loops to graphs: engineering multi-agent systems

A loop lets one agent correct itself. A graph lets many loops check one another, and grounding keeps all of them tied to what is actually true.

Following the buzz around loops engineering, I have started to see another term turn up beside it: graphs engineering. I'm not sure the label will stick. Names in this field have a short half-life, and most of them describe a product category before they describe an idea. The shift underneath this one feels real to me, though, and it sits close enough to the work I do every day that I wanted to think it through properly rather than in a few lines of a post.

A loop is a good unit of work

Most early agent systems were built around a single loop. The agent plans, acts, checks the result and tries again, until the check passes or the budget runs out.

A single agent loopA task starts the loop. The agent plans, acts, and checks the result; if the check fails it revises the plan and goes round again; if it passes, the loop ends. starts executes returns result fails updates passes Task Plan Act Check Revise Done one agent one objective, one check AGENT STEP CHECK
Fig. 1 One agent, one objective, one check. The loop ends when the check passes.

The loop deserves more credit than the new vocabulary gives it. It is the reason coding agents became useful at all: an agent that can run the tests, read the failure and try again will beat a cleverer one that answers once and hopes. The idea is also much older than agents. In 1960 Miller, Galanter and Pribram described behaviour in terms of TOTE units (test, operate, test, exit) and argued that plans are built from loops like these, nested inside one another.

A loop has one limit built into it, which is that it is only as good as its check, and its check is usually local. An agent fixing a failing test knows when the test passes. It doesn't know whether the test was the right one, whether the fix broke something another team depends on, or whether someone should have been asked first. For a contained task that is fine, because the task is the whole world. Inside a real business process a single loop stops being enough very quickly, because the process is not one task with one check but many tasks, owned by different people, with checks that belong to someone else.

What a real process looks like

Due diligence is one of the workflows I build agents for, so it is the example I reach for. The graph below shows the typical shape of the process, the review a buyer runs over a company's documents before a deal, rather than any particular system.

A due diligence process as a graph of agentsA data room supplies a triage agent, which routes documents to four review streams running in parallel: contracts, corporate, employment, and IP and regulatory. An entity map shares identities with the streams. The streams pass evidence to a verifier, which can reopen a single stream. Confirmed issues go to a risk register, which escalates to a lawyer who can stop the work. The report waits on both the register and the lawyer's approval. supplies routes resolves parties shares identities passes evidence reopens one stream confirms escalates feeds approves Data room Triage Entity map Contracts Corporate Employment IP & regulatory Verifier Risk register Lawyer review can stop the work Report waits on both RUNS IN PARALLEL AGENT CHECK KNOWLEDGE PERSON RERUN
Fig. 2 A typical due diligence process. Review streams run in parallel; the verifier can reopen one stream without restarting the others; the report waits for both the risk register and a lawyer's sign-off.

Different agents handle different parts of the work. One triages the data room and routes each document to the right stream. Several review streams run in parallel, for commercial contracts, corporate records, employment and intellectual property, while an entity map keeps track of which company is which. A verifier checks what the streams extracted against the source documents. Issues collect in a risk register, a lawyer reviews whatever has been flagged, and the report is drafted only when both are done.

Two things in that picture have no place in a loop. Some steps run at the same time, because reviewing employment contracts doesn't depend on reviewing leases. Others have to wait, because a decision upstream changes what they should be doing: if the parties turn out to include a subsidiary nobody listed, every stream needs to know before it finishes. The work has the shape of a dependency graph whether anyone draws it or not, and a system that doesn't know that shape will either run everything in sequence or let some streams finish on facts that have already changed.

The interesting part is the edges

When people share diagrams like this one, the conversation tends to be about how many agents are in it, which I find the least interesting thing about them. What matters is how they work together, and nearly every question I care about turns out to be about the edges rather than the nodes.

Start with checking. A check is only worth something if it is able to disagree with what it checks, and a verifier that reads the reviewer's summary, with the same model and the same style of prompt, will mostly agree with it. A verifier earns its place by going back to the source, or by using a different method, and ideally both.

Stopping is a different kind of edge, because it needs authority as well as ability. In legal work a missing document or a possible conflict should halt the review rather than be logged and passed along, and someone has to own that decision. In the systems I build, that someone is a lawyer. Human-in-the-loop review is there so that the lawyer keeps the judgement, and guardrails decide what each agent is allowed to see and do.

Disagreement is where I think most designs go wrong. When two agents reach different answers, the tempting fixes are to average them, to take a vote, or to let whichever ran last win, and all three throw away the most useful signal the system produces. A disagreement should travel, with both positions and the evidence behind each, to whoever is entitled to settle it. Often that is a lawyer; occasionally it is a rule written down in advance.

The last question sounds like plumbing. When one check fails, does a small part of the workflow run again, or does everything start over? Software settled a version of this in 1976, when Stuart Feldman wrote Make, which rebuilds only the targets whose dependencies have changed because it knows the graph. An agent system that knows its own graph can rerun the contracts stream after one extraction fails verification and leave the other streams alone. On a large deal, that decides whether the system is usable at all.

None of these questions are new. Every firm already answers them about its people: who drafts, who reviews, who signs, who can say no. Separation of duties and the four-eyes principle are graph design for organisations. What has changed is that we now have to answer them explicitly, in code, for workers who will never stop to ask.

Building agents for legal work showed me that the model is the replaceable part. I keep the models swappable on purpose, and the engineering lives in the harness around them: the tools, the workflow, the review and the data you feed in. A graph is the same idea one level up, a harness for a group of agents that decides what each one sees, what it may do, and who checks its work.

A graph can be wrong in a very organised way

There is an obvious catch. A graph of agents looks like a system of checks, and it can still be wrong in a very organised way. If every agent relies on the same incomplete data, or every check is scored by the same weak metric, they can all agree and still miss reality.

Agreement without ground, and the grounded versionLeft: an extractor, reviewer and verifier all read the same data room, which is missing a side letter; all three agree there is no issue. Right: claims cite their sources, an entity graph says which documents should exist, and a gap check flags the missing side letter to a lawyer. reads agree reads cites sources confirms rereads the source expects documents lists contents flags what is missing Data room Side letter never uploaded Extractor Reviewer Verifier No issue found three agreements, one gap Data room Extractor Verifier Finding with sources Entity graph what should exist Gap check Lawyer asks for the letter AGREEMENT WITHOUT GROUND GROUNDED
Fig. 3 Left, three agents agree because they all read the same incomplete data room. Right, claims cite their sources, and a gap check that knows what should exist notices the letter that never arrived.

Take the due diligence case. Suppose a side letter that changes the termination terms of a key contract never made it into the data room. The reviewer reads the contract and finds no issue. The verifier checks the reviewer's summary against the contract and confirms it. The risk register stays clean, and the report goes out with three layers of agreement behind a conclusion that is wrong. Every agent did its job, and the graph did exactly what it was designed to do.

Condorcet described this failure in 1785, although he was writing about juries. His jury theorem says that a majority is more likely to be right than any single voter, but only when each voter is better than chance and they make their mistakes independently. Remove the independence and every extra voter adds confidence without adding accuracy. Agents built on the same model, reading the same documents and scored by the same metric are about as far from independent as voters can get.

So agreement inside the graph is not evidence on its own. Something has to connect it to the world outside.

Grounding is the second graph

That something is grounding, and it is the part of this shift I care about most, because it is where my own work has always been. Before I built agents, I spent years building knowledge graphs in the life sciences, and the step that mattered most was rarely the clever one. It was grounding: resolving every mention of a gene, a compound or a disease to one identifier in a shared ontology before anything else touched it. Without that, two systems could agree about "the same" protein while talking about two different things. Drug discovery taught me a second version of the same lesson. A prediction from the graph earned a scientist's trust only when it came with an evidence chain, a path through the graph they could read link by link and reject link by link if one of them was wrong.

Agent graphs need both lessons, and in practice they come down to three commitments. Edges should carry evidence and not only conclusions, so that when one agent hands a finding to another it also passes pointers to the sources it relied on, and the next node can go back to the document instead of trusting the summary. The entities the agents talk about, the parties, the contracts, the obligations, should resolve to one shared model of the domain, so that "the target" means the same company everywhere in the graph. And the system should know what it expects to see, because if a contract refers to a side letter, the absence of that letter is a finding in its own right.

That last commitment is the one that catches the organised error, and a pile of documents can't meet it. It needs a model of the domain, of what exists and how it relates, and that model is itself a graph. This is why, alongside the agents, I'm designing a knowledge graph of the firm's own legal knowledge, so that agents reason over it rather than over scattered documents. A serious agent system ends up with two graphs: a workflow graph that says who does what and who checks whom, and a knowledge graph that says what is there, how it connects, and where each claim came from. I often put this as moat = workflow + knowledge, and the two graphs are what that formula looks like when you draw it.

People stay in the picture on purpose. Grounding keeps the system connected to the evidence, while people remain responsible for the judgements that matter. A good graph doesn't ask a lawyer to reread everything the agents read. It brings them the decisions that need their judgement, with the evidence attached and any disagreement visible, and it gives them the authority to stop the work.

From prompt to grounded organisation

As I see it now, the progression looks something like this.

Prompt → Chain → Loop → Graph → Grounded organisation

From prompt to grounded organisationFive stages from left to right: a single prompt; a chain of three steps; a loop of three steps; a graph of several connected loops; and the same graph resting on a shared knowledge layer, with a person in it. Each stage contains the one before. Prompt ASKS ONCE Chain FIXED STEPS Loop CHECKS ITSELF Graph LOOPS CHECK EACH OTHER Grounded organisation SHARED GROUND, CLEAR OWNERS
Fig. 4 Each stage contains the last. Graphs are full of loops, and loops are full of chains.

Each stage keeps the one before it. A prompt asks once, a chain fixes the order of several prompts, and a loop lets one step check its own work and try again. A graph runs many loops side by side and lets them pass evidence along and correct one another. A grounded organisation is a graph whose agents and people stand on the same ground, with the same sources, the same model of the domain, and clear lines of who is accountable for what.

I'm less sure about the last step than the others. It is partly a description of where the strongest systems are heading and partly a description of how good firms already work, and I don't yet know how much of an organisation can be made explicit enough for agents to share it. That is the question I find myself designing around.

Whether or not graphs engineering survives as a name, I think that is the real work. We are not replacing loops. We are designing the larger system around them, and giving it ground to stand on.

Jiakang Chang · Principal Software EngineerAll articles