A pilot has just succeeded. The new system, service or process was tried at a couple of sites, the measured result beat the target, the people involved are enthusiastic, and the sponsor is asking for money to roll it out everywhere. The case feels stronger than most business cases, because it rests on something that actually happened rather than a forecast.
Before rolling it out, look closely at what the evidence can and cannot show. A pilot usually proves that something can work. That is valuable, and it is a smaller claim than most pilot reports make. It does not, on its own, show that the result will hold without the extra support the pilot enjoyed, at the price the rollout assumes, more than once, or with people who did not volunteer. A pilot with all those advantages and one without them can produce the same headline number, and the difference lies in conditions nobody wrote down.
This article explains the three questions a pilot result should answer before it justifies a larger commitment, why pilots are usually populated by the most willing people, how to design a pilot that tells you more, and what to do when a rollout disappoints. It is general information for owners and managers who trial new ways of working before committing to them.
Three questions the result must answer
Ask of any pilot result:
- Would it have happened without the extra help? Pilots attract unusual inputs: the best staff, the owner’s attention, a supplier’s goodwill, rules bent to get things moving, someone fixing data problems after hours. None of this is improper. All of it is support the full rollout will not receive.
- Would it work at the price the rollout assumes? If the pilot site got the service free, or the cost was absorbed centrally, or a long-standing customer accepted a discount as a favour, the most important assumption in the business case, that someone will pay this price, has not been tested.
- Has it happened more than once? One successful cycle shows a capable team, given attention, can do it once. A process is something that produces the result again with different people, at a different site, without the original team present.
A pilot that cannot answer all three is not worthless. It should be recorded as promising, not proven, and the next step should be designed to answer the missing questions.
Selection happens twice
Most pilots are populated by people who wanted them, and that happens in two places:
- Design. The people who help shape the new system are usually those who are available and supportive. The design fits them.
- Trial. Pilot sites are usually volunteers. That is sensible, because volunteers find and report problems. But it means the trial population was chosen by its enthusiasm.
Volunteers differ from everyone else in predictable ways:
| How they differ | Typical effect |
|---|---|
| Setup | Newer equipment, cleaner data, fewer local quirks |
| Capacity | A team volunteers when it has someone to spare; busy teams do not |
| Location | Central or well-connected sites volunteer more readily than remote ones |
| Incentive | Volunteers often gain from the change; others may lose something |
| Skill | Pilot teams tend to include the best people, hiding how much the design depends on skill |
A pilot with willing participants answers “does it work?” well. It cannot answer “will everyone else use it?”, because the people who would not have chosen it were never in it.
One caution. The fact that some sites did not volunteer says nothing about their motives. They may have been short-staffed, mid-audit or already running something that works. An objection raised later by a group never asked earlier is information about the design, not evidence of resistance.
Design a pilot that tells you more
A few changes make a pilot far more informative:
- Record conditions as you go, not just outcomes. Who was involved, what extra support was given, what was expedited, what was free, what rules were relaxed. A result reported without its conditions cannot be interpreted later.
- Write a coverage statement before the pilot starts: which kinds of site, team or customer are included, which are not, how the absent ones differ, and what evidence you would need before the absence stops mattering.
- Include at least one site that would not have volunteered. Not a hostile one; a representative one. It makes the pilot harder and the results less flattering, and it is usually the most valuable change you can make.
- Add a withdrawal period. After the result is achieved, remove the extra inputs, such as the owner’s visits, the extra staff and the waived costs, and keep measuring. What survives is the evidence. What disappears was the support, and knowing its size is worth knowing.
- Repeat in stages. First the same team at a different site, which tests local conditions. Then a different team at a different site, which tests dependence on particular people. Somewhere in that sequence, the sponsor steps back.
- Test the price. If the rollout assumes users pay a fee or a department funds it, have the pilot site pay something close to that price, at least for a period.
- Compare with a baseline. Results during a pilot can be flattered by seasonal changes, a quiet month or other improvements happening at the same time. Compare with similar sites that did not take part, or with the same period last year. The what would have happened anyway article covers setting a realistic base case.
- Report feasibility and adoption separately. Say plainly what the pilot showed about whether the thing works, and what it can and cannot show about whether it will be used.
Use the pilot to find problems, not just to prove success
Enthusiastic pilot teams are often the best testers a business has. They push a new system harder, report defects instead of quietly working around them and suggest fixes. That makes the pilot’s most valuable output not the headline number but the list of everything that had to be adapted, fixed or worked around to get there. Ask the pilot team to keep that list as they go: the manual steps, the exceptions, the data that had to be cleaned, the questions staff asked repeatedly. Each item is a cost or a risk the rollout will meet at every other site.
Someone pays for the transition
Adopting anything new costs the adopters something: time to learn, disruption while old and new overlap, and sometimes a loss of familiar discretion. In a pilot, willing participants absorb that cost quietly, often with extra help. At scale, it falls on teams that did not ask for the change and may receive less of its benefit. Budget for that transition support explicitly, and decide who pays for it, rather than assuming the pilot’s goodwill will be repeated everywhere.
Write the scale-up trigger in advance
Before the pilot result exists, write down the condition for rolling out: the result, under what conditions, at what cost, repeated how many times, and who decides. A trigger written afterwards tends to be a rationalisation of whatever happened, and one set by the sponsor alone is not a control. The learn before you commit article covers timing tests so the results arrive before the point of no return.
When a rollout disappoints, diagnose before you add rigour
If a rollout underperforms, the usual response is more governance: tighter reporting, a stronger project lead, a recovery plan. That is the right response when the idea was sound and delivery was poor. It is the wrong response for other kinds of failure:
| What you see at scale | What has usually failed | What more rigour achieves |
|---|---|---|
| Adoption only where the sponsor is present | The evidence: the pilot measured sponsorship | A bigger dependence on sponsors |
| Heavy discounting needed to sell it | Positioning: customers cannot tell it apart | Better-managed discounting |
| Volume grows but profit does not | Sequencing: expansion came before proof | Efficient expansion of a loss |
| Delivery late against a sound design | Execution | The intended fix |
Identify which layer failed before deciding what to do. The delivered is not adopted article covers making a change stick after rollout.
A worked example
This is an illustration. A commercial cleaning business with 15 sites trials a new rostering and shift app at two sites. The supervisors at both sites volunteered, the operations manager visits each weekly, the app supplier waives the subscription during the trial, and the office manager spends evenings cleaning up the staff data. After eight weeks, the time supervisors spend on rostering falls from about 10 hours a week per site to about 6, a 40% saving, and no-shows fall.
The sponsor proposes rolling out to all 15 sites. Before agreeing, the owner asks the three questions and writes a coverage statement.
- Extra help: weekly visits, evening data work and two keen supervisors.
- Price: the app was free; at scale it costs about $8 per user a month, around $14,400 a year for 150 staff.
- Repetition: once, at two similar day-shift sites.
- Absent from the pilot: night-shift-only sites, sites whose supervisors are less confident with English, and staff with older phones.
The owner adds a third stage. The operations manager stops visiting the two pilot sites for six weeks while measurement continues; the saving there settles at about 30%. The app is also introduced at a representative night-shift site that would not have volunteered, without extra support. There, the saving is only about 15%: supervisors have little quiet time to set up rosters, and some staff struggle with the app’s language settings.
The business case is redone with the lower figure. At an illustrative $40 an hour, a 15% saving across 15 sites, about 1.5 hours a site a week over 48 weeks, is worth about $43,200 a year, comfortably above the $14,400 subscription. The original pilot figure would have suggested about $115,200. The rollout goes ahead, but phased, with a translated quick guide, a half-day setup session for each night-shift supervisor and a support budget, and the owner’s expectations are set by the representative site rather than the volunteers.
How this applies to a small Australian business
- Ask the three questions: without the extra help, at the real price, more than once.
- Record pilot conditions alongside the results.
- Write a coverage statement of who is in and out of the pilot.
- Include a site that would not have volunteered.
- Remove the extra support and keep measuring.
- Test the price the rollout assumes.
- Write the rollout trigger before the result arrives.
- Diagnose disappointing rollouts before adding more control.
Signals worth watching
- Pilot reports with outcomes but no conditions.
- Pilots run only at the most enthusiastic sites.
- A free trial used to justify a paid rollout.
- One successful run treated as proof.
- Adoption that stalls wherever the sponsor is absent.
- Non-volunteering sites labelled as resistant.
- Pilot workarounds that never made it into the rollout plan.
Common mistakes
- Treating “it worked” as “it will be used”.
- Forgetting the extra help the pilot received.
- Never testing the price.
- Scaling after a single success.
- Dismissing objections from groups never asked.
- Responding to every rollout problem with more governance.
Frequently asked questions
Should we stop using volunteer sites? No. Volunteers are good at finding problems. Just do not read their results as a forecast for everyone, and add a representative site.
Isn’t removing support unfair to the pilot team? Frame it as protecting their result from being misread. It shows how much of the success will carry over.
How many repetitions are enough? Usually two or three under progressively normal conditions, depending on the size of the commitment.
What if we cannot afford a longer pilot? Then you are choosing speed over certainty, which can be reasonable. Record that choice and budget for the risk.
Does this apply to trialling a new product with customers? Yes. A product bought mainly by friends, loyal customers or people given a special price shows it can be delivered, not that the wider market will pay for it. Test with customers who have no reason to do you a favour.
Who should decide whether a pilot has proven enough? Someone other than its sponsor, using criteria written before the result was known.
Questions to ask
- What extra help did this pilot get, and what did it cost?
- What price did the pilot site pay, and what will the rollout assume?
- How many times, and under what conditions, has the result been achieved?
- Who was not in the pilot, and how do they differ?
- What happened when the extra support was removed?
- What condition, written in advance, would justify rolling out?
Bringing it together
A successful pilot shows that something can work. Before it justifies a larger commitment, check that the result survives without the extra help, at the real price and more than once, and remember that volunteers are not a sample of everyone else. Record conditions as well as outcomes, include a site that would not have volunteered, remove the support and keep measuring, and write the rollout trigger before the result arrives. When a rollout disappoints, find which layer failed before adding more control.
Source: KEVOS notes, drawing on teaching material on pilot evaluation and transformation evidence, including a 2005 Purchasing trade-journal account of a development bank’s procurement system rollout using volunteer departments. Examples and figures in this article are illustrations. This article is general information.