A consultant will charge you $5,000 to $10,000 to tell you where AI fits in your business. Sometimes that's money well spent. Often it isn't, because the first pass at this question doesn't really need expertise. It needs somebody inside the business paying attention for two weeks.
What follows is the front half of the readiness work we do with clients. Not a summary of it, the actual thing: the logging, the scoring, the cut lines, and the part where the answer comes back "not yet."
You don't need to know anything about AI to do this. You need to know where your week goes.
One rule before you start
Log what happens, not what's supposed to happen.
Every business has a documented process and a real one, and they're never quite the same. The documented one lives in a Google Doc somebody wrote two years ago. The real one lives in a person's head and involves a spreadsheet they built themselves, a folder called FINAL_v3, and a text message to whoever's on site.
Automate the documented process and you'll build something nobody uses. So for the next two weeks you're writing down the real one, including the embarrassing parts.
Week 1, days 1–3: the time log
Pick the three to five people whose time you most wish you had more of. Usually that's you, whoever runs operations, and whoever answers the phone. For three working days, each of them logs every task in four columns:
- What they did. One line, plain language. "Retyped job details from the voicemail into the scheduler."
- How long it took. Real minutes, not the number that sounds reasonable.
- How often it happens. Daily, weekly, monthly.
- Judgment or steps. Did they have to weigh something up, or just follow a sequence?
Two things go wrong at this stage. People round down, because writing "40 minutes" next to a task feels like a confession. And they leave out the small stuff, which is where most of the time actually hides: a five-minute task done twelve times a day is bigger than the two-hour task done once a week.
Three days is enough. You're not building a complete record, you're after the shape of the week.
Days 4–5: sort into three buckets
Now go back through the log and put every task in exactly one bucket.
Repetitive
Same inputs, same steps, same kind of output. The test I use: could you write instructions a new hire could follow on their second day? If yes, it's repetitive, and this is the only bucket that moves forward.
Judgment-heavy
Needs context, tradeoffs, a read on the situation. AI can draft a first version of these and often should. It shouldn't decide them, which means the payoff is smaller and the risk is higher. Come back to these later.
Relationship-critical
The value here is that a person did it. The renewal call with your biggest client. The apology after something went wrong. Automating these saves twenty minutes and costs you the account.
Week 2: score what's left
Take the repetitive bucket and give each task four numbers.
- Hours per month. Frequency × duration, added up across everyone who touches it.
- Error cost, 0–3. What happens when it's done wrong? Zero means nobody notices. Three means you lose money or a customer.
- Data readiness, 0–3. Is the information this task needs somewhere a system can reach it? Zero means someone's head or a filing cabinet. Three means a system with an export button.
- Change rate, high or low. How often does the process itself change? Anything that changes monthly will need somebody maintaining whatever you put on top of it.
Score is hours per month × (error cost + data readiness). Then knock anything with a high change rate down a tier.
There's nothing magic about the formula. It exists to make you compare tasks on the same axes instead of ranking them by whichever one annoyed you most this morning. Data readiness is the number people fudge, and I'd push back hardest on that one. It's the best predictor of whether a build goes smoothly or turns into a six-week data cleanup nobody budgeted for.
The cut line
Sort by score and draw two lines.
- Under 10 hours a month: leave it. Scoping it will cost more than it returns. Write it down, look again in six months.
- 10 to 40 hours a month: buy or configure something that already exists. Don't build. There is almost certainly a tool for this, and your job is to pick one rather than invent one.
- Over 40 hours a month, with data readiness of 2 or 3: worth building properly, with an owner, a measurement, and a plan for month six.
A typical audit at a 40-person company produces one item in the top tier, two or three in the middle, and a long tail below the line. If yours produces eight in the top tier, you scored too generously. Go back and be meaner about data readiness.
What people usually find
Different industries, same three things near the top of the list almost every time.
- Intake. The first contact, whether that's a call, a form, or an email, gets handled by whoever's free. So it gets handled inconsistently and sometimes late. This is the most common top score we see and the one owners most often price at zero.
- Follow-up. The second, third, and fourth touch that nobody has time for. Quotes that were never chased. Invoices thirty days past due that nobody called about.
- Reporting. Someone spends a day a month pulling numbers out of four systems into a deck that gets skimmed once.
If your audit surfaces something that looks nothing like these three, check the scoring first. But do look at it properly, because it's either a mis-scored task or the most interesting thing on your list.
When the answer is "not yet"
Sometimes two weeks of honest logging tells you not to start. The data lives in three systems that don't talk to each other. The process changes every month because the business is still working out how it wants to run. The one person who understands the whole thing is leaving in March.
That's a real result and a cheap one. Gartner expects more than 40% of agentic AI projects to be canceled by 2027, and the reasons they give are unclear value, rising costs, and weak governance rather than anything to do with the technology. Two weeks of logging is the cheapest possible version of checking first.
When to stop doing it yourself
The audit tells you where. It won't tell you how, or what it costs, or which vendor is overselling. Three points where outside help earns its money:
- Vendor selection. Forty tools claim to do the same thing, and comparing them takes your own test cases plus somebody who knows which questions end a bad sales call early.
- The top-tier build. If something scored above the line, the difference between a system that's working in month six and one that quietly rotted is product management, not code.
- Disagreement on the scores. Two people scoring the same task wildly differently usually means the process isn't understood as well as everyone assumed. Worth an outside read before you build anything on top of it.
Finish the two weeks holding a ranked list and a clear top three and you've done the part most businesses never get to. That list is worth more than any vendor demo, and it's yours whether or not you ever hire anyone.
If two weeks is more than you can spare right now, our AI Readiness Scorecard takes about fifteen minutes and covers the same axes at lower resolution. And if the audit tells you you're stuck between too many options rather than short of them, that's a different problem with a different fix, which we wrote about in the AI decision gap.