jkchang.uk
Article · 2 Sep 2026Agentic systems · Context engineering9 min read

From tickets to behavioural contracts: specifying work for AI agents

Tickets split work into pieces a person can hold. Once agents do the execution, the scarce thing is a precise statement of what the system should do, and that statement becomes the agent's brief, its test and the audit record.

Whether a team called itself Waterfall or Agile, the work tended to reach engineers in the same form. A product manager or business analyst read the requirements, broke them into epics, broke the epics into stories and the stories into tasks, and handed the smallest pieces to the people who would build them. The ticket at the bottom of that tree is a familiar object: a title, a line of the form "As a user, I want..., so that...", three or four acceptance criteria, an estimate and somebody's name.

Tickets did their job for a long time, and my doubt is about who they were designed for. They assume a particular kind of executor, and as coding agents take on more of the execution, the difficult part of the work moves to a place tickets were never built to hold. It moves from breaking the work down to being precise about how the system should behave. I'm increasingly convinced that the way we organise software work has to follow it there.

Tickets were designed around human limits

The case for decomposition is strong, and it deserves a fair hearing before I argue with it. People need bounded context, because nobody can hold a whole system in their head, and a task that fits into a day or two can be understood, estimated and finished by one person. People also need clear ownership, since work with no name on it tends not to get done, and manageable units are what let a team see progress and notice when something is stuck. Fred Brooks put the cost of ignoring this into arithmetic in 1975: a team of n people has n(n−1)/2 possible lines of communication, so every person added makes coordination more expensive. A good breakdown keeps most of those lines quiet. Each engineer has to understand their own piece and the edges where it meets the next one, and not much else.

So the model made sense while humans were the main execution layer, and it still makes sense for coordinating humans. What I had underestimated is how much of a ticket's success depended on things that were never written in it.

Ron Jeffries described the user story in 2001 as three parts: the card, the conversation and the confirmation. The card was meant to be small, a reminder rather than a specification. The real requirement lived in the conversation between the people who wanted the feature and the people building it, and the confirmation was the set of tests that showed it was done. In practice that conversation carried on long after the planning meeting. An engineer who hit an ambiguity could ask the product manager, check with whoever built the neighbouring feature, or simply know, from two years on the team, what the business would want. The ticket was an index into a large body of shared human context, and it worked because that context was there to be looked up.

Agents move the bottleneck

That shared context is the part that doesn't transfer. A coding agent such as Claude Code or GitHub Copilot's coding agent can read the requirement, inspect the codebase, change files, run the tests and iterate, often faster than I can review what it did. What it can't do is walk over and ask what the product manager meant. When a ticket is ambiguous, the agent resolves the ambiguity itself, usually silently and often plausibly, from the repository and from whatever the model assumes a typical product would want. Give it a ticket written for a human and it will close the ticket. Whether it built the behaviour you needed depends on how much of your intent survived the trip into a few lines of text.

Meanwhile the decomposition itself has become cheap. Breaking a requirement into a list of tasks is something current models do competently, and the spec-driven tools that appeared in 2025 make that explicit. AWS's Kiro turns a requirement into a requirements document, a design and a task list, and GitHub's Spec Kit has separate steps for specifying, planning and generating tasks. In both, the task list is generated output, and the human input everything else depends on is the specification at the top.

So the bottleneck moves. When decomposition is cheap and execution is fast, the scarce thing is a clear, checkable statement of how the system should behave: what a user should be able to do, what must never happen, and how anyone would know the work is finished. I usually describe an agent as model + harness, and most of what I have written about harnesses has been about tools, context and boundaries (see Less role play, more boundaries). The specification belongs in the same list. Of everything in the harness, it is the only part that says what the work is for.

Behaviour is the unit an agent can check

This is why I'm moving towards a workflow borrowed from behaviour-driven development. The steps are easy to state: write the user requirement directly, define the specific system behaviour it needs, let the agent build against that behaviour, and then use the completed BDD documentation as the material for validation, audit and review.

BDD is about twenty years old. Dan North introduced it in 2006, after years of teaching test-driven development, because he found that the word "test" kept sending people in the wrong direction, while writing each test as a sentence about what the system should do made it easier to decide what to test first and what to call it. With Chris Matts, a business analyst, he developed the Given/When/Then template that later became the Gherkin syntax used by Cucumber. The aim from the start was shared understanding between the people who want a system and the people who build it, written in a form that could also be run.

Here is an invented example, to show the shape rather than any real system:

Requirement
  A user can search every document they are allowed to see,
  and learn nothing about the documents they are not.

Scenario: restricted documents stay out of search
  Given a user who can read the "Public" folder but not the "Board" folder
  And both folders contain a document that mentions "merger"
  When the user searches for "merger"
  Then the results contain the document from "Public"
  And the result count is 1
  And nothing in the response refers to the "Board" document

Boundaries
  Permissions are applied before ranking and counting.
  The permissions schema does not change.

Evidence
  The scenario passes in CI.
  A second test changes the user's permissions between two searches.

Each part does a job that a ticket either does loosely or leaves to the conversation. The requirement says what matters to the user. The scenario says what success looks like, in a form the agent can run against while it works instead of discovering at review. The boundaries say what can't be crossed, which a ticket rarely states because an engineer on the team would already know it. The evidence says what has to exist before anyone calls the work complete. The line about the result count is the kind of detail I mean. A ticket that said "hide restricted documents from search" could be satisfied by code that filtered the list after ranking and left the total at 2, which tells the user that a second document exists.

I use the word contract on purpose. Bertrand Meyer's design by contract, which he built into the Eiffel language in the late 1980s, describes a routine by its preconditions, its postconditions and the invariants it must preserve, so that what a piece of code promises can be stated separately from how it keeps the promise. A behavioural specification for an agent does the same job one level up. The Givens are preconditions, the Thens are postconditions, and the boundaries are invariants: things that must still hold however the agent chooses to implement the change.

BDD's record with human teams is mixed, and the reason matters here. A common complaint about Cucumber suites is that the business stopped reading the feature files, so engineers were left maintaining two layers, the English scenarios and the step definitions that made them run, for an audience that had gone away. Scenarios decayed into brittle scripts of clicks. My guess is that agents change that calculation in two ways. The specification now has a reader that reads all of it every time, because the agent works from it directly. And the glue code that made BDD tedious is the kind of work agents are good at, so the cost of keeping scenarios executable has dropped. I hold that guess loosely, because I am still early in working this way.

Documentation becomes the contract

In this model, documentation is no longer just a planning artefact. It becomes the execution contract, telling the agent what matters, which boundaries it cannot cross, what success looks like, and what evidence is needed to prove the work is complete.

That reverses a long habit. The Agile Manifesto in 2001 valued "working software over comprehensive documentation", and with good reason: most specification documents of the time were written once, read rarely and wrong within months, because nothing forced them to stay true. A behavioural contract is under a different kind of pressure. The agent builds from it and the scenarios run against it, so when the code and the contract disagree, something fails. Documentation that is executed stays honest in a way that documentation which is only read never did.

The completed contract is also the most useful thing to give a reviewer. Reviewing an agent's pull request cold, from the diff alone, has the weakness I wrote about in From loops to graphs, where a verifier checks the summary when it should check the source. A reviewer who has the contract can ask better questions. Does each scenario describe what the user actually needed? Is a boundary missing? Does the evidence show the behaviour, or only that some tests passed? The review moves up a level, from reading code line by line to checking behaviour against intent, and the line-by-line review that remains has something to check against.

Audit is the part I care about most, because of where I work. In legal practice "show me why this is right" is an ordinary question, and the agent systems I build there have to give answers that can be defended. I think the software that produces those answers should meet the same standard. When someone asks, months later, why a system behaves the way it does, a ticket history can say who did what and when. A behavioural contract, with the evidence that it was met, says what the system was required to do and how anyone knew it did. That is the record I would want to hand an auditor.

Where the contract breaks

I don't want to oversell this, because the failure modes are real. The first is that a contract can be satisfied to the letter. Models trained to make tests pass sometimes take the cheap route, and Anthropic's system card for Claude 3.7 Sonnet reported the model special-casing tests instead of fixing the underlying problem. A scenario is a test with better prose, and it can be gamed the same way. The obvious defences are to have a person review the scenarios before the agent writes any code, and to include in the evidence checks that the agent did not write, so it isn't marking its own homework.

The second is that writing behaviour precisely is hard, and much of it is exactly the thinking a vague ticket let us postpone. Teams used to find the edge cases while building. Now someone has to find them before the agent starts, or accept that the agent will meet them first and decide on its own. That cost is real, and it lands on whoever owns the requirement. The third is that not everything worth specifying fits a scenario. Performance and accessibility can be written as constraints, but some of what makes software good is still judgement that only shows up in review.

Tickets don't disappear either. People still need to know who owns what and when it will land, and a task list is still a good way to coordinate the humans around an agent. What changes is which document is the source of truth. The ticket tracks the work, and the contract defines it.

What senior people write now

I think this changes what product and engineering leadership spend their time on. Much of the craft used to be decomposition: taking a large, vague goal and cutting it into pieces a team could pick up without stepping on each other. AI can now generate the tasks, write the code and propose implementation paths, and it will keep getting better at all three. It can't decide what the system should do for the people who use it, which risks matter, or what will need to be auditable a year from now. Those decisions need someone who understands the users, the domain and the consequences, and the behavioural contract is where they get written down.

Jeffries's story card was always a promise of a conversation. The agent doing the work was never part of that conversation and can't join it, so the conversation has to be written down, precisely enough to build from and to check against afterwards. Writing it well has become, to my mind, the most senior work left in the process, and no amount of better tickets will stand in for it.

Jiakang Chang · Principal Software EngineerAll articles