All posts
Operating model

The first 90 days of an AI agent

Launch is the easy part. Here's what breaks in week 1, week 3, week 6, and week 12, and the fifteen hours of attention that decide whether it's better or worse at day 90.

Published September 29, 2026
Reading time 7 min
By Dominic Morgan

Every AI deployment I've helped launch got a launch plan. Almost none of them got a plan for the ninety days after — which is odd, because that's the stretch that decides whether the thing was worth buying.

The useful part is that those ninety days are predictable. The same four things go wrong at roughly the same four moments, in businesses with nothing else in common. You can put them in the calendar before you go live.

Week 1: the volume shock

Testing gave the agent a curated set of inputs, one at a time, in order. Launch gives it Tuesday morning. Eleven things at once, the person who calls three times in an hour, the email that says "?", the caller who starts mid-sentence because they thought they were still on hold.

The visible symptom is an escalation rate two or three times what the test set predicted. Alarming and mostly fine. Week-one escalations are largely the agent doing what it should — being careful in situations it hasn't seen before. An agent that escalates too much on day three is a fixable problem. An agent that guesses confidently is a much worse one.

So do nothing. Watch, log, categorize at the end of the week, and judge no metric on week-one numbers.

I've made the opposite mistake myself. Day two of an agent I'd shipped, escalation count spiked, and I widened its authority to bring the number down. Which worked, in the sense that the number went down. It also hid a real problem for a week — the agent was now confidently handling exactly the cases it had been right to escalate. That's the shape of every week-two panic move I've seen since.

Week 3: the edge cases arrive

By week three the ordinary work is handled and the long tail shows up. The customer who's also a supplier. The job type you stopped offering last year that's still on page four of the website. The person who answers a question with a question. The request that's technically in scope and obviously shouldn't be.

You now have three weeks of escalations, which is the first genuinely useful dataset the project has produced. Sort it by frequency, not by which one annoyed somebody most — because that's how it'll get sorted if nobody insists otherwise.

Fix the top three categories and deliberately ignore the rest. A tail addressed one case at a time produces a set of instructions so long and so full of exceptions that the agent gets worse at the common work. Rare things are supposed to escalate. That's the design working.

This is also the first real change to the system, so it goes through the gate rather than around it. Run the new version against the stored results of the launch version, and have somebody read the outputs that changed most before it goes out. If that doesn't mean anything concrete to you yet, this is the piece that explains it. Expect to remove capability here about as often as you add it.

Week 6: boredom

Nothing is broken. Nobody is complaining. The novelty has worn off, and the person who checked the dashboard every morning for a fortnight hasn't opened it in nine days.

This is the dangerous one. It's the exact point where a project and a product go their separate ways. Nothing about week six feels like it needs attention, which is the property that makes it the moment systems start dying. I've seen more agents die of week-six neglect than of any single failure mode.

The only defense that survives a busy quarter is a recurring calendar entry and a report that arrives whether or not anybody asks. Thirty minutes: skim the week's escalations, look at two numbers, decide whether anything needs doing. Usually nothing does, and that's a result worth having in writing.

Week six is also the first honest moment to ask what would make this more useful, because now the answer comes from six weeks of real usage instead of what everyone assumed during scoping. It's almost always different from what everyone assumed during scoping.

Week 12: the drift nobody noticed

Nothing about the agent has changed. Everything around it has.

Prices went up in October. You stopped covering one postcode. The warranty terms changed. Somebody left and the handoff goes to a different person now. Each of those is a fact the agent states confidently, in the same steady voice it uses for the facts that are still true.

This is the most common way I've seen an AI system actually die, and it happens somewhere between month three and month six. Not a crash. A slow accumulation of small wrongnesses that nobody catches until a customer does.

Two defenses, both cheap. The first is a quarterly fact review — write down every fact the agent states as true (prices, hours, service area, policies, names, promises) and put a date and an owner against each. Forty-five minutes once a quarter, and most lines take ten seconds. The second is rerunning the test set: the suite that gated the launch is the only thing that catches drift before a customer has to. Facts that changed show up as new failures on examples that used to pass.

Week twelve is also when the expansion question gets a real answer. There's usually pressure to add a second job by now, generally from whoever is happiest with the first one. Expand only if the core job is genuinely solid. Bolting a second capability onto a shaky first one reliably produces two mediocre ones and an owner who can no longer tell which part is underperforming.

What ninety days of this costs

Week one is about two hours, spent watching rather than changing. Weeks two through four are an hour a week, mostly reading escalations, plus one prompt change. Weeks five through twelve are an hour a month, plus a forty-five-minute quarterly review.

Under fifteen hours across the whole quarter. That's the real price of a system that's better at day 90 than at launch, and it's why the work needs a named owner. Fifteen hours is nothing. Fifteen hours belonging to nobody is fifteen hours that don't happen.

The version where nobody does it

Worth spelling out, because it isn't dramatic and that's precisely why it's common.

Escalations pile up unread. The instructions never change, so the top three gaps from week three are still there in month five. Drift accumulates. Usage decays a little each month as people who got one stale answer go back to doing it themselves without mentioning it. Nothing ever breaks loudly enough to trigger a response.

Then at month nine somebody asks whether this is worth renewing, and nobody in the room can answer. The system gets canceled — not because it failed but because no one can prove it worked.

Which is the case for treating these as products rather than projects, made at more length in products, not projects. The ninety-day calendar above is what that argument looks like once you put dates on it.

Want to talk about this in your business?

Most of what we write comes out of actual client work. If anything here resonates, we'd be happy to dig into the specifics with you.