The pilot that dies without anyone killing it

You know the shape of it. Someone brings an AI tool to a Tuesday meeting. It looks useful. Everyone agrees it is worth evaluating, and a couple of people get assigned to look into it.

It becomes a standing agenda item. Six weeks in, two people have logged in once. Twelve weeks in, the question has grown: should we compare three vendors instead of one? Should somebody review the data policy? By month five the original champion has a different priority, and the item quietly drops off the agenda.

Nobody made a decision. Nobody killed the project. It ran out of oxygen.

In a company of 12 or 30 people, this is the most common outcome of an AI initiative. Not a failure and not a success. An expiration.

Long evaluations are not more careful. They are just longer.

The theory behind a six month evaluation is that time reduces risk. More vendors compared, more stakeholders consulted, more edge cases mapped, better decision at the end.

That theory holds when the thing being evaluated sits still. AI tooling does not sit still. The product you shortlisted in March has shipped several releases by September, changed its pricing tiers, and added the feature that was your main objection. Two competitors you never looked at now do the same job. The evaluation is not accumulating accuracy. It is accumulating staleness.

There is a second cost that is easier to miss. A small business runs on a fixed attention budget, and an open evaluation spends it whether or not anything happens. Every week the pilot stays open is a week the owner is carrying it in the back of their head. That carry is real, and it is not free.

What the failure data actually says

Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027. The three reasons it names are escalating costs, unclear business value, and inadequate risk controls. Two of those three are scoping failures, not technology failures. A project with a number attached to it does not have unclear business value.

You have probably also seen the MIT NANDA figure that 95% of generative AI pilots return nothing measurable. Treat that one with some suspicion. It rests on 52 interviews and 153 survey responses, all self reported, and the authors describe their own numbers as directionally accurate rather than precise. But one finding buried inside it is worth more than the headline: deployments done with an outside partner reached production roughly twice as often as ones built internally.

That gap is not evidence that outside help is smarter. It is evidence that an engagement has a start date, an end date, and someone whose job is to finish.

A pilot with no end date is not a pilot. It is a subscription to thinking about it.

Meanwhile the gap between company sizes is widening. Census Bureau data through early May 2026 puts AI use at 37% among firms with at least 250 employees, while firms under 20 people sat below 20% and stayed close to flat over the prior six months. Companies your size are not losing an AI arms race. Most of them never entered it, and an open ended evaluation is one of the more respectable ways to not enter.

The two week version

Replace the evaluation with a commitment that has edges. Four things, decided before anyone logs in to anything.

One workflow. Not "AI for customer service." One named, repeated, annoying task: inbound quote requests, appointment confirmations, intake triage, invoice coding, the weekly report somebody rebuilds by hand. It should happen at least a few times a week so you get signal quickly.

One owner. A person, not a committee, and preferably the person who does the work today rather than the most technical person in the building. If nobody will put their name on it, you have learned something useful already and it cost you nothing.

One number. What will be different, and measured how. Hours per week on the task. Days to first response. Share of quotes followed up inside 24 hours. Pick the number you would actually be willing to report out loud.

One deadline. Two weeks for most things. Long enough to do the work honestly, short enough that it cannot quietly become furniture.

Then the step most teams skip. Write down in advance what result makes you keep it and what result makes you stop. A pilot with no stated kill condition does not end, it fades, which is failing slowly and learning nothing on the way.

Two weeks is enough more often than you would guess

The obvious objection is fair: two weeks is not enough to know. For a platform decision, that is correct. For the first useful thing, it usually is not.

Most high value automation in a small business is narrow. It touches one form, one inbox, one spreadsheet, one recurring handoff between two people. The hard part is almost never model selection. It is writing down how the process actually works today, which most companies have never done, and deciding what the software is allowed to do without asking a human first. Two weeks is a reasonable amount of time for that, and running it is the only way to find out.

If your honest answer is that nobody has ever written the process down, start there. That is a prerequisite, not a delay, and it is worth its own short post: your business is only as big as what you wrote down.

Shipping once also produces something no evaluation ever does: a working reference point. After one automation is live you know how your team reacts to it, where your data is messier than you assumed, and what "good" costs in real hours. Every decision after that gets cheaper. That is the compounding asset, and research cannot buy it.

When a slow evaluation is the right call

Sometimes it is. If the workflow touches regulated data, if your client contracts carry data handling terms, if a mistake would not be reversible, then slow is correct. Be slow deliberately, with a written scope and a date on the calendar. That is a decision with a timeline, which is a different thing entirely.

What does not qualify is still comparing vendors in month four because nobody wants to be the person who picked wrong. That is risk aversion wearing diligence as a costume. It carries its own risk, which is that another year goes by and the manual process is still the process. If the hesitation is really about governance rather than tooling, that is an operations conversation, and it is the kind of thing Highland Private Office exists to sort out.

The honest summary

The main risk for a 5 to 50 person business is not picking the wrong AI tool. Tools are cheap and switchable, and a bad one costs you a month. The real risk is spending two quarters deciding, shipping nothing, and then concluding that AI does not work for a company like yours, when what did not work was the evaluation.

Pick one workflow. Give it an owner, a number, and two weeks. Decide up front what would make you stop. At the end you either have a working thing or a clear answer, and both of those beat a standing agenda item.