jkchang.uk
Article · 2 Jul 2026Agentic systems · Context engineering7 min read

The third generation of skills

Skills began as prompts and grew scripts. The next ones will carry judgement instead, and leave the script to the agent.

When Anthropic published its Claude for Legal plugins in May, I read through the repository expecting to find the usual anatomy of an agent skill: a page of instructions, a few reference files, and a folder of Python doing the parts that shouldn't be left to the model. The repository holds twelve plugins and 151 skills, for commercial, corporate, employment, IP, litigation, privacy and regulatory work among others, and none of those skills ships a script. The only code sits at the top level, where it validates and deploys the plugins themselves; everything inside the skills is written guidance. At first that read to me like a step backwards, until it started to look like the next step.

Three generations, briefly

The first generation of skills was mostly prompts. A skill was a textual workflow: instructions, rules, a few examples and a description of the expected behaviour, packaged so that an agent could load it when the task called for it. It worked well enough for tasks where the model's own judgement was sufficient, and badly for anything that needed the same precise result every time.

The second generation added scripts. When Anthropic introduced Agent Skills in October 2025, it described a skill as a folder of instructions, scripts and resources, and its main worked example was a PDF skill that carries a Python script for extracting a form's fields. The reasoning was sound. Some operations are cheaper, faster and more reliable as ordinary code than as generated text, and nobody wants a model improvising a PDF parser token by token every time a form arrives. Prompts plus repeatable actions made skills much more dependable, because not everything should be left to language generation.

The third generation, as I see it, keeps the guidance and drops most of the pre-packaged code. The skill defines the principles, the boundaries, the success criteria and the safety rules, and the agent writes whatever script the task in front of it needs. That is what the legal plugins look like, and I don't think it's an accident of one repository.

Three generations of skillsFirst generation: a prompt instructs the agent, which produces an output. Second generation: the prompt instructs the agent and a script bundled in advance fixes part of the path. Third generation: guidance sets the boundaries, the agent writes and runs a script for the task in front of it, a check tests the result against the success criteria, and the agent stops to ask a person when a boundary is reached. instructs instructs fixes the path bounds writes and runs stops and asks Prompt Agent Output Prompt Bundled script written in advance Agent Output Guidance boundaries, criteria Agent Script written for this task Criteria met? Person Prompts FIRST GENERATION Prompts + scripts SECOND GENERATION Guidance THIRD GENERATION
Fig. 1 Three generations of skills. In the second, a script written in advance fixes part of the path; in the third, guidance sets the boundaries and the agent writes the script for the task in front of it.

The case for scripts is real

Before arguing the other way, the second generation deserves a fair hearing, because its advantages are the ones a legal team should care about most. A script is deterministic. It does the same thing on Tuesday as on Monday, and when it fails, it fails loudly with a stack trace instead of quietly with a plausible paragraph. It can be tested, reviewed and versioned like any other code. It costs almost nothing to run compared with a model reasoning its way to the same result. In a field where an output may have to be defended to a client or a court, all of that matters, and you would expect legal work to be the last place to give it up. So I wanted to know why a suite built for exactly that field doesn't contain a single script.

A script is one context, frozen

The trouble with a bundled script is that it encodes decisions made for the situation its author had in mind. It assumes a document layout, a naming convention, a sequence of steps. When a user calls the skill, the agent follows that predetermined path, and the path holds only as far as the author could see.

Real inputs rarely stay inside that foresight. The agreement arrives as a scan, the amendments are named final.pdf and final-v2.pdf, one clause has been renumbered twice. Each edge case becomes another branch in the script, or another special-case instruction telling the agent when not to run it. As the cases pile up, the skill turns into a decision tree that is hard to maintain and still misses the next case, while the agent, which could have read the documents and worked out what to do, sits idle following someone else's plan.

Turn it around. If the agent decides for itself and writes the script it needs for this matter, these documents and this request, the same skill can cover far more situations and handle edge cases as they come, instead of after someone has patched them in. The script stops being an asset you maintain and becomes a by-product of one run, which is what it always was: one context's answer. What the agent cannot work out for itself is what the script is for, and what it must never do.

The runtime has moved as well

The platforms have been moving the same way as the skill files. Anthropic launched Claude Managed Agents in public beta on 8 April 2026: an Anthropic-run harness with sandboxing, tool execution, state and scoped permissions, which you hand a task and get a result back from. OpenAI's Agents API takes the same shape. A request carries a model, instructions and an input, and the agent runs in an OpenAI-hosted sandbox. The quickstart's example task asks the agent to create tree.py, a script that prints the files in the current directory, then run it and show the output. Its instructions read, in full, "Write clean code, run it, and report the actual output."

That example is small, but it captures the change. The caller supplies a task and a standard; the agent supplies the code. When every agent comes with a sandbox to write and run code in, a pre-written script inside a skill starts to constrain the agent more than it protects the user, and the scarce thing becomes knowing what code should be written, and when none should be.

What is worth writing down

If the script is disposable, the valuable part of a skill is the human judgement behind it. As far as I can tell, that judgement comes down to five questions:

  1. When should we automate?
  2. What needs to be verified?
  3. What should never be touched?
  4. When should the agent stop and ask?
  5. What does "done" actually mean?

The legal plugins answer all five, and seeing them answered in a real skill is more persuasive than the list. Take the amendment-history skill, which traces how a contract has changed across its base agreement and every amendment. On automation, it tells the agent to infer from the request whether the user wants a summary of all changes or a trace of one clause, and not to ask which mode is wanted unless the request is genuinely ambiguous. On verification, it insists on establishing the chronological order of the documents before reading their content, using execution dates where they exist, dates in the recitals where they don't, and the references each amendment makes to the agreement it modifies to confirm the chain.

On what not to touch, it forbids reading another matter's files unless cross-matter context has been switched on. On stopping to ask, it names the conditions precisely: filenames that give no sequence, dates missing from both filenames and headers, two documents that look like the same amendment. Outside those cases the agent proceeds, and where it inferred the order rather than confirmed it, it says so at the top of its output. On done, the output carries the work-product header and inherits the privilege status of its sources.

The rules around one run of the amendment-history skillA request arrives and the agent picks a mode, asking only if the request is ambiguous. It fixes the chronological order of the documents from dates, recitals and cross-references, and asks the user only when no sequence can be found. It then reads and traces the changes and produces an output that carries the work-product header. All of this happens inside the active matter; other matters' files stay off limits. arrives summary or trace confirmed chain no sequence found confirms Request Pick the mode Fix the order Read and trace Output Ask the user only when there is no sequence Other matters never read INSIDE THE ACTIVE MATTER ask only if ambiguous dates, recitals, references in chronological order work-product header AGENT CHECK KNOWLEDGE PERSON
Fig. 2 The amendment-history skill drawn as the rules around one run. None of it is code: each step is a decision the skill's authors wrote down, and the agent works out the rest.

None of this is code, and all of it is what a careful associate would need telling in their first week. A script could implement any one of these rules for one document set. The rules are what survive to the next one.

Guidance is harder to write than code

The third generation is not free, because it moves the difficulty from programming into description, and description has its own failure mode. An agent working from guidance will often find a good solution. When it doesn't, the cause is more often the guidance than the agent: the problem wasn't described clearly enough, or the wording quietly pushed the agent towards a decision it shouldn't have made. A line written with one case in mind, say "always ask the user to confirm ordering", gets applied to every case, and an agent that would have reasoned well is steered into reasoning badly. A broken script fails where you can see it. Misleading guidance fails inside a confident answer.

So the discipline is to describe the problem rather than the answer: what the documents are, what can go wrong with them, what must hold at the end, and where the edges are. The amendment-history skill does this with its ordering rule. It doesn't tell the agent to confirm the order every time; it tells it to ask only when the filenames and dates give no sequence, which leaves the judgement with the agent and still marks where the judgement runs out.

It also explains something I first found odd about the repository. Its skills aren't short. There are about 35,000 lines of Markdown across them, and the longest, an employment skill for internal investigations, runs to 765. Length was never the measure for either generation, though. The best future skills won't necessarily be the longest ones, or the ones with the most scripts. They will be the clearest, where every line is a decision someone made about the work, written so the agent can't mistake it for a decision about something else.

A good third-generation skill gives the agent room to act and keeps the boundaries firm: flexible execution, clear constraints, verifiable outcomes. The repository I read through contains 151 skills and not one script. I expect that to look ordinary quite soon.

Jiakang Chang · Principal Software EngineerAll articles