Ask for a list of every agent touching production systems and you will usually get three answers that disagree: the one the platform team maintains, the one implied by vendor spend in the finance ledger, and the one an executive last presented to the board. Reconciling those three lists is the first thing a new agent owner inherits. It is not the interesting part of the mandate, but it decides whether the remaining sixty days produce anything shippable or sixty days of maintenance on work nobody would authorise today.
The kill list is the output of that reconciliation. It is a named, dated document stating which pilots stop, which are archived with their assets preserved, and which continue with a business owner attached. Published in week four, it buys back capacity. Left unpublished, the estate grows for another quarter and every subsequent cut costs more political capital than this one would have.
What the inherited estate usually contains
Three failure patterns account for most of what a new owner finds. They look different in a status report and are killed for different reasons.
- Pilots with no named business owner. Built by a capable engineer or an external partner, sponsored by a function that has since reorganised. The line in the roadmap says “in production”. Nobody in the operating business has the outcome in their objectives, so nobody notices when quality drifts.
- Agents holding standing write access nobody reviews. A service account created for a two-week trial, still able to write to a CRM, a ticketing system or a payments ledger months later. The pilot may have stopped producing value long ago; the credential did not stop working.
- Demos kept alive because a sponsor is attached to them. These are the expensive ones. The work is real, the sponsor is senior, and the demo has appeared in a town hall or an investor deck. It has never processed a live transaction without a human rewriting the output.
A fourth category deserves separation: pilots that are genuinely early and genuinely on track. They also lack hard numbers, and an audit run carelessly kills them along with the rest. The criteria below are written to distinguish the two.
The audit questions, one page per pilot
Run the same questionnaire against every item, including the ones you built or championed yourself. Deviating from a fixed set is how exceptions enter the process. One page per pilot, answers sourced from systems rather than from the person who owns the pilot, is enough.
- Which named individual, at director level or above, carries the outcome of this pilot in their objectives this year?
- What decision or transaction does the agent’s output feed, and what happened to that decision before the agent existed?
- How many times did it run in the last 30 days, and what share of those runs produced output a human accepted without material edits?
- What baseline was measured before go-live, by whom, and where is it recorded?
- Which systems can it write to, under which credential, and when was that credential last reviewed?
- What is the monthly run cost, split into model or API spend, licences, and internal engineering hours?
- Who is on call when it produces a wrong output, and what is the documented rollback?
- If it were switched off this afternoon, who would notice within a week, and how?
The last question does most of the work. A pilot whose disappearance nobody would detect inside seven days is not in production regardless of how it is labelled in the portfolio.
Termination criteria and the evidence that settles them
Publish the criteria before publishing the verdicts. A pilot owner who sees the rule first and the judgement second argues about the rule; one who sees the judgement first argues about the person who made it. Thresholds below are policy settings, not industry benchmarks – set them to your own risk tolerance and state the numbers you chose in writing.
| Criterion | Evidence that settles it | Default disposition |
|---|---|---|
| No named business owner at director level or above | Objectives document or performance plan naming the outcome | Stop within 10 working days unless an owner accepts it in writing |
| No measured pre-launch baseline, and none reconstructable | Dated baseline record with method and sample size | Stop, or restart as a scoped 6-week experiment with a baseline first |
| Human edits or overrides the majority of outputs | Acceptance logs over a minimum 30-day window | Stop the agent, keep the evaluation set and the prompts |
| Standing write access to a system of record with no review record | Access review with a signed reviewer and date | Revoke write access within 48 hours, downgrade to read-only pending review |
| Run cost exceeds the manual cost of the same work | Monthly cost breakdown against loaded hourly cost of the displaced task | Stop unless a costed path to parity exists within one quarter |
| Duplicate of another pilot in scope and data | Side-by-side scope comparison | Merge into the stronger instance, stop the other |
| Demo status for more than two quarters with no live transaction | Transaction log or its absence | Archive with assets preserved, reopen only against a named use case |
Two guards keep this from becoming a purge. First, a pilot inside its first 90 days with a written baseline and a named owner is exempt from the outcome criteria; it is too early and the audit says so explicitly. Second, anything with a regulatory or safety dependency goes to the relevant control function before it is stopped, because switching off an agent that produces an audit trail is itself a change requiring approval.
Access revocation as a separate track
Stopping a pilot and revoking its credentials are different pieces of work, and the second is routinely forgotten. A decommissioned agent whose service account retains write access to the CRM is worse than a running one, because nothing is watching its output any more.
Run the access track ahead of the kill list and independently of it. Enumerate every service account, API key and OAuth grant issued to an agent or its orchestration layer. Record the scope, the issuing owner and the last use. Anything with write access to a system of record and no named reviewer goes to read-only immediately, whatever the pilot’s fate. This is not a negotiation with the pilot owner; it is standard access hygiene applied to a class of accounts that skipped it during the pilot rush.
Announcing a shutdown without losing the sponsor
The sponsor of a killed demo is often the person whose support you need for the work that replaces it. The announcement sequence matters more than its wording.
- Brief the sponsor before anyone else, in person, with the criteria in hand. Not the verdict alone – the rule, the evidence against it, and the fact that the same rule was applied to every item including your own.
- Attribute the decision to the portfolio, not to the pilot’s quality. The honest framing is usually true: the estate cannot support fourteen half-owned pilots, so seven stop and three get real funding.
- Preserve and name the assets. Evaluation sets, labelled data, prompt libraries and integration work survive the shutdown and get reused. Say where they are stored and who owns them.
- Offer a re-entry route with a stated condition. “This reopens when the operations director puts the metric in their objectives” is a route. “We may revisit later” is not, and sponsors read it correctly as a refusal.
- Give one of the sponsor’s people a role in what continues. A sponsor with a seat at the surviving work rarely litigates the shutdown.
Publish the full list at once rather than closing pilots one at a time. Serial cuts create four separate arguments and a queue of people lobbying not to be next. A single dated document with the criteria attached creates one argument, held once.
Sequence across the first four weeks
- Week one: assemble the inventory from three independent sources – platform logs, finance spend and the reported roadmap. Treat the disagreement between them as a finding worth reporting on its own.
- Week two: run the audit questionnaire, one page per pilot, with answers pulled from systems. Open the access enumeration in parallel and revoke unreviewed write access as it is found.
- Week three: circulate the criteria, not the verdicts. Give owners five working days to submit evidence you could not find. Some will produce a baseline that exists but was never filed.
- Week four: brief sponsors individually, then publish the list with dispositions, dates, asset locations and re-entry conditions.
Budget the shutdowns as work. Decommissioning has a cost: credential revocation, data retention decisions, contract termination notice, and telling any external party that the engagement ends. Pilots that were cheap to start are not always cheap to stop, and a kill list without an owner and a delivery date for each shutdown is a memo, not a decision.
The call in week four
By the end of the first month the evidence is either sufficient or it is not, and both outcomes point to the same choice. Publish the kill list with names, dates and criteria attached, absorb the objections in a single week, and enter day thirty-one with a smaller estate and free engineering capacity. Or keep the full inventory running, and spend the next sixty days maintaining work that would not survive a funding review today – while the pilots you would actually start queue behind it.
Write the list this week. Put your own favourite pilot on it if it fails the criteria, and lead the briefing with that one.