From loops to graphs: engineering multi-agent systems
A loop lets one agent correct itself. A graph lets many loops check one another, and grounding keeps all of them tied to what is actually true.
Following the buzz around loops engineering, I have started to see another term turn up beside it: graphs engineering. I'm not sure the label will stick. Names in this field have a short half-life, and most of them describe a product category before they describe an idea. The shift underneath this one feels real to me, though, and it sits close enough to the work I do every day that I wanted to think it through properly rather than in a few lines of a post.
A loop is a good unit of work
Most early agent systems were built around a single loop. The agent plans, acts, checks the result and tries again, until the check passes or the budget runs out.
The loop deserves more credit than the new vocabulary gives it. It is the reason coding agents became useful at all: an agent that can run the tests, read the failure and try again will beat a cleverer one that answers once and hopes. The idea is also much older than agents. In 1960 Miller, Galanter and Pribram described behaviour in terms of TOTE units (test, operate, test, exit) and argued that plans are built from loops like these, nested inside one another.
A loop has one limit built into it, which is that it is only as good as its check, and its check is usually local. An agent fixing a failing test knows when the test passes. It doesn't know whether the test was the right one, whether the fix broke something another team depends on, or whether someone should have been asked first. For a contained task that is fine, because the task is the whole world. Inside a real business process a single loop stops being enough very quickly, because the process is not one task with one check but many tasks, owned by different people, with checks that belong to someone else.
What a real process looks like
Due diligence is one of the workflows I build agents for, so it is the example I reach for. The graph below shows the typical shape of the process, the review a buyer runs over a company's documents before a deal, rather than any particular system.
Different agents handle different parts of the work. One triages the data room and routes each document to the right stream. Several review streams run in parallel, for commercial contracts, corporate records, employment and intellectual property, while an entity map keeps track of which company is which. A verifier checks what the streams extracted against the source documents. Issues collect in a risk register, a lawyer reviews whatever has been flagged, and the report is drafted only when both are done.
Two things in that picture have no place in a loop. Some steps run at the same time, because reviewing employment contracts doesn't depend on reviewing leases. Others have to wait, because a decision upstream changes what they should be doing: if the parties turn out to include a subsidiary nobody listed, every stream needs to know before it finishes. The work has the shape of a dependency graph whether anyone draws it or not, and a system that doesn't know that shape will either run everything in sequence or let some streams finish on facts that have already changed.
The interesting part is the edges
When people share diagrams like this one, the conversation tends to be about how many agents are in it, which I find the least interesting thing about them. What matters is how they work together, and nearly every question I care about turns out to be about the edges rather than the nodes.
Start with checking. A check is only worth something if it is able to disagree with what it checks, and a verifier that reads the reviewer's summary, with the same model and the same style of prompt, will mostly agree with it. A verifier earns its place by going back to the source, or by using a different method, and ideally both.
Stopping is a different kind of edge, because it needs authority as well as ability. In legal work a missing document or a possible conflict should halt the review rather than be logged and passed along, and someone has to own that decision. In the systems I build, that someone is a lawyer. Human-in-the-loop review is there so that the lawyer keeps the judgement, and guardrails decide what each agent is allowed to see and do.
Disagreement is where I think most designs go wrong. When two agents reach different answers, the tempting fixes are to average them, to take a vote, or to let whichever ran last win, and all three throw away the most useful signal the system produces. A disagreement should travel, with both positions and the evidence behind each, to whoever is entitled to settle it. Often that is a lawyer; occasionally it is a rule written down in advance.
The last question sounds like plumbing. When one check fails, does a small part of the workflow run again, or does everything start over? Software settled a version of this in 1976, when Stuart Feldman wrote Make, which rebuilds only the targets whose dependencies have changed because it knows the graph. An agent system that knows its own graph can rerun the contracts stream after one extraction fails verification and leave the other streams alone. On a large deal, that decides whether the system is usable at all.
None of these questions are new. Every firm already answers them about its people: who drafts, who reviews, who signs, who can say no. Separation of duties and the four-eyes principle are graph design for organisations. What has changed is that we now have to answer them explicitly, in code, for workers who will never stop to ask.
Building agents for legal work showed me that the model is the replaceable part. I keep the models swappable on purpose, and the engineering lives in the harness around them: the tools, the workflow, the review and the data you feed in. A graph is the same idea one level up, a harness for a group of agents that decides what each one sees, what it may do, and who checks its work.
A graph can be wrong in a very organised way
There is an obvious catch. A graph of agents looks like a system of checks, and it can still be wrong in a very organised way. If every agent relies on the same incomplete data, or every check is scored by the same weak metric, they can all agree and still miss reality.
Take the due diligence case. Suppose a side letter that changes the termination terms of a key contract never made it into the data room. The reviewer reads the contract and finds no issue. The verifier checks the reviewer's summary against the contract and confirms it. The risk register stays clean, and the report goes out with three layers of agreement behind a conclusion that is wrong. Every agent did its job, and the graph did exactly what it was designed to do.
Condorcet described this failure in 1785, although he was writing about juries. His jury theorem says that a majority is more likely to be right than any single voter, but only when each voter is better than chance and they make their mistakes independently. Remove the independence and every extra voter adds confidence without adding accuracy. Agents built on the same model, reading the same documents and scored by the same metric are about as far from independent as voters can get.
So agreement inside the graph is not evidence on its own. Something has to connect it to the world outside.
Grounding is the second graph
That something is grounding, and it is the part of this shift I care about most, because it is where my own work has always been. Before I built agents, I spent years building knowledge graphs in the life sciences, and the step that mattered most was rarely the clever one. It was grounding: resolving every mention of a gene, a compound or a disease to one identifier in a shared ontology before anything else touched it. Without that, two systems could agree about "the same" protein while talking about two different things. Drug discovery taught me a second version of the same lesson. A prediction from the graph earned a scientist's trust only when it came with an evidence chain, a path through the graph they could read link by link and reject link by link if one of them was wrong.
Agent graphs need both lessons, and in practice they come down to three commitments. Edges should carry evidence and not only conclusions, so that when one agent hands a finding to another it also passes pointers to the sources it relied on, and the next node can go back to the document instead of trusting the summary. The entities the agents talk about, the parties, the contracts, the obligations, should resolve to one shared model of the domain, so that "the target" means the same company everywhere in the graph. And the system should know what it expects to see, because if a contract refers to a side letter, the absence of that letter is a finding in its own right.
That last commitment is the one that catches the organised error, and a pile of documents can't meet it. It needs a model of the domain, of what exists and how it relates, and that model is itself a graph. This is why, alongside the agents, I'm designing a knowledge graph of the firm's own legal knowledge, so that agents reason over it rather than over scattered documents. A serious agent system ends up with two graphs: a workflow graph that says who does what and who checks whom, and a knowledge graph that says what is there, how it connects, and where each claim came from. I often put this as moat = workflow + knowledge, and the two graphs are what that formula looks like when you draw it.
People stay in the picture on purpose. Grounding keeps the system connected to the evidence, while people remain responsible for the judgements that matter. A good graph doesn't ask a lawyer to reread everything the agents read. It brings them the decisions that need their judgement, with the evidence attached and any disagreement visible, and it gives them the authority to stop the work.
From prompt to grounded organisation
As I see it now, the progression looks something like this.
Prompt → Chain → Loop → Graph → Grounded organisation
Each stage keeps the one before it. A prompt asks once, a chain fixes the order of several prompts, and a loop lets one step check its own work and try again. A graph runs many loops side by side and lets them pass evidence along and correct one another. A grounded organisation is a graph whose agents and people stand on the same ground, with the same sources, the same model of the domain, and clear lines of who is accountable for what.
I'm less sure about the last step than the others. It is partly a description of where the strongest systems are heading and partly a description of how good firms already work, and I don't yet know how much of an organisation can be made explicit enough for agents to share it. That is the question I find myself designing around.
Whether or not graphs engineering survives as a name, I think that is the real work. We are not replacing loops. We are designing the larger system around them, and giving it ground to stand on.