M.A.I. Consulting
GuideAdvanced1 October 20264 min read

Named owner, operating guide, test set: what makes an agent survive its author

An agent only its builder understands is a personal project running in production. A named owner, a written operating guide and a periodic test set are the minimum bar before calling it finished.

This closes the cluster this series has spent five pieces building: what an agent actually is, why the method has to come first, and where the absolute limits sit. This piece answers the practical question that follows once an agent is genuinely built and working: what stops it quietly breaking, or quietly drifting into something nobody understands, the day the person who built it moves on. The honest answer is three artefacts, none of them technical, all of them cheap compared to the cost of finding out the hard way that nobody else can touch the thing.

An agent that only its builder understands is not actually finished. It is a personal tool wearing the appearance of organisational infrastructure, and the gap between those two things only becomes visible at the worst possible moment, usually right after the builder has already left.

A named owner, distinct from the builder

The person who built the agent and the person who owns it going forward do not need to be the same person, and for most organisations this size, keeping them the same is actually the risk. Ownership here means something specific: one named person, not a team, who is accountable for the agent still doing what it is supposed to do, who gets told first when something looks wrong, and who has the authority to switch it off. If the builder leaves and nobody was ever named as owner, the agent does not fail cleanly. It keeps running, unattended, until something breaks badly enough that someone finally asks who is meant to be watching it.

An operating guide the owner did not have to write themselves

The operating guide is the artefact that actually transfers the knowledge the builder is carrying around in their head: what the agent does, what it does not do, what normal output looks like, what a wrong output looks like, and what to do when one shows up. This is not a technical specification. It is closer to the method this series argued for building before the agent existed in the first place, now updated to describe what actually got built rather than what was intended. An owner who inherits an agent with no operating guide is not really an owner. They are a person with responsibility for something they cannot actually evaluate.

A test set that catches drift before a person does

The test set is a small, fixed collection of real inputs with known-correct outputs, run periodically against the agent to confirm it still behaves the way it did when it was accepted. This matters because the way an agent quietly stops working is rarely a dramatic failure. It is a slow drift, a platform update, a small change in the documents it reads, a shift in the task itself, that nobody notices until the output has been subtly wrong for weeks. A test set turns that invisible drift into something a person can actually check on a schedule, rather than something that only surfaces once a beneficiary or a funder points it out.

What each artefact actually protects against

A working agent with no named owner, no operating guide, and no test set is not really an organisational asset. It is a personal project that happens to be running in production, and the organisation will not find out the difference until the person who built it is no longer there to explain it.

The practical takeaway

Treat a named owner, a written operating guide, and a periodic test set as the minimum bar for calling an agent finished, not optional extras added once time allows, since without all three the agent only survives for as long as its builder stays.

No external statistic cited; this article presents an internal practical framework rather than third-party evidence.

Series · Deciding what to automate · part 4 of 4
Keep reading
05 ยท Custom Agents & Tools

Can the tool we already pay for do this task for us, every time?

If this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.