← Back to Insights
6 min read
To run an AI pilot that pays for itself in 60 days, pick one painful, high-volume process, define the money it wastes before you start, build the smallest automation that moves that number, and measure the result against the baseline. Keep the scope brutally narrow, put a person in the loop, and decide to ship or stop on evidence - not on how clever the demo looks.
I run pilots like this for a living, so I will be blunt about the part most guides skip: the pilot is not the technology, it is the proof. If you cannot say in advance what "it worked" means in pounds and hours, you are not running a pilot - you are funding an experiment with no finish line.
A working pilot produces a nice demo. A paying pilot produces a number your finance lead recognises. The difference is that you fixed the target before you built anything: the hours a process burns, the cost of the errors it makes, or the revenue it leaks. This matters because most AI projects never clear that bar. Gartner projected that at least 30% of generative AI projects will be abandoned after proof of concept, often because they were never tied to a measurable business outcome in the first place (Gartner). A pilot that pays for itself is simply one you designed to be measured.
Because it is long enough to build something real and short enough to force discipline. A 60-day box kills the two ways pilots usually fail: the endless "let's add one more thing" scope creep, and the science project that never has to prove anything. It also matches how the value tends to land - most of the operational wins I deliver come from a few weeks of focused work on one process, not a six-month platform programme. The clock is a feature. It stops you from confusing motion with progress.
Choose a single workflow that is high-volume, rule-heavy and genuinely painful - the invoice that gets keyed in by hand fifty times a week, the enquiry that waits a day for a human to triage it. Resist the urge to pick the most exciting use case; pick the most expensive boring one. And make sure it is actually ready: a broken process automated is just a faster broken process. If you are not sure, run it through this readiness test before you commit a single day to it.
This is the step that separates a pilot from a punt, and it takes an afternoon, not a fortnight. Measure the process as it runs today:
Write it down and get your sponsor to agree it. If you skip this, you will have no honest way to prove the pilot worked - and "it feels faster" convinces no board.
Automate one step, not the whole department. The goal is the thinnest slice that touches the metric you chose: classify the incoming email, extract the fields off the invoice, draft the reply for a human to approve. Build it where your data already lives rather than moving everything into a new tool - integrating into your existing stack is faster and stickier than a rebuild, as I argued in build vs buy vs integrate. Keep the model swappable behind a thin layer so you are not welding the pilot to one vendor before you even know it works.
For anything customer-facing or financial, the pilot should draft and a person should approve. This is not a lack of ambition - it is how you build trust and catch failure modes cheaply while the stakes are low. It also gives you a clean measurement: you can see exactly how often the AI got it right before you ever consider letting it run unattended. Autonomy is something you earn with evidence, not something you switch on because the demo looked confident.
Point the pilot at live volume - or a fair sample of it - for a couple of weeks, and compare the result to the baseline you wrote in Step 2. Not to your hopes; to the number. Did turnaround drop? Did the error rate fall? How many drafts did a human have to correct? Real inputs surface the messy edge cases a demo never will, and that mess is exactly what you need to see before you trust it.
At the end of the box, you make an honest call against the target:
All three are wins, because all three are decisions backed by evidence rather than vibes. The only failure is a pilot that limps on indefinitely, proving nothing.
Less than the process it fixes, if you scope it properly. A tightly bounded pilot is exactly the kind of fixed-scope, proof-of-value work that should cost thousands, not tens of thousands - and it should be sized so the annual saving dwarfs the build. The maths only works because the scope is narrow: one process, one metric, one clear decision at the end. Widen the scope and you lose the economics; hold the line and the pilot funds the next phase itself.
A pilot pays for itself when you treat it as a measurement exercise, not a technology showcase. Pick one expensive process, write down what it costs you today, build the smallest thing that moves that number, keep a person in the loop, and decide to ship or stop on the evidence. Do that inside 60 days and you either walk away with a banked saving or a cheap, early "no" - both of which beat a demo that impresses everyone and changes nothing.
That is exactly how I run first engagements. If you have a process bleeding time and want a proof-of-value pilot scoped to pay for itself, that is the kind of thing a thirty-minute discovery call sorts out - we pick the process, agree the number, and I build the smallest thing that moves it. See process automation and AI consulting for how I approach it.