Skip to content
Skip to content

Scope a Pilot That Can Actually Finish

Most AI pilots do not fail on the technology. They fail because nobody wrote down what success looked like, who owned it, or when it would stop. This checklist covers the decisions to settle before the first line of work starts.

Work through it with the person who will own the pilot. If an item cannot be ticked, that is the conversation to have before spending money.

AI Pilot Scoping Checklist

Twenty-four decisions to settle before the pilot starts.

0 of 24 complete0%

Your ticks are saved in this browser, so you can work through the list over several sessions.

01Scope and ownership

0/5

One workflow, one person, and a written list of what is excluded.

02Measurement

0/5

A number before, a threshold, and the conditions for stopping.

03Data and access

0/5

Confirmed with a real sample, not assumed from a system list.

04Human review and safety

0/4

Recommend-only first. A person checks before anything leaves.

05Budget, time and communication

0/5

A cap, an end date, and the people who need to know.

A practical scoping list, not a project management method. Items on data handling are general guidance and not legal advice. Your ticks are stored in this browser only and nothing is transmitted.

Why Scoping Decides the Outcome

A pilot is an experiment. An experiment without a measured baseline, a threshold and an end date cannot produce a result, only opinions.

One workflow, one owner

A pilot that covers three processes and reports to a committee produces three half-answers and no decision. Narrow the scope to a single workflow with a single named person who can say yes or no at the end.

A number before, a number after

If you do not measure the workflow before the pilot starts, you cannot show what changed. The baseline is the cheapest and most skipped step, and it is the one that makes the final decision defensible.

An end date and an exit

Pilots without a time box drift into permanent trial. Set the end date at the start, and write down what happens at that date in each of the three cases: scale it, fix it, or stop.

The Five Areas

Twenty-four items across five areas. Each item is a decision you can make in a meeting, not a document you need to write.

1

Scope and ownership

The one workflow in scope, the one person accountable, and the list of what is deliberately excluded.

2

Measurement

The baseline metric measured today, the threshold that counts as success, and the exit criteria for stopping early.

3

Data and access

Confirmation that the data and systems the pilot needs are reachable, with permission, before work starts.

4

Human review and safety

Who checks the output, what the system may never do unattended, and who can switch it off.

5

Budget, time and communication

A capped spend, a fixed end date, and a plan for telling the people the pilot affects.

The Items Most Often Skipped

Four items from the list that are missing from most pilot plans we see, and what happens when they are.

The baseline metric

Pick one number that describes the workflow today: minutes per item, items per day, error rate, backlog size, or days to respond. Measure it for at least two normal weeks before the pilot touches anything. Without this, the end-of-pilot review becomes a debate about impressions.

  • One metric, measured the same way before and after
  • Two normal weeks of baseline, not a quiet week or a peak week
  • Record who measured it and how, so it can be repeated
  • If it cannot be measured, the workflow is not ready for a pilot

Data access confirmed, not assumed

The most common reason a pilot stalls in week two is that the data turned out to be in a system with no export, a licence tier without API access, or a folder nobody has permission to read. Confirm access with a real sample before the start date.

  • Pull a real sample of the data the pilot will use
  • Check the licence tier of every system involved for API or export access
  • Confirm who approves access and how long that takes
  • Note any personal information and how it will be handled

A human review path

A pilot should start in recommend-only mode. The system drafts, suggests or classifies, and a named person reviews before anything reaches a customer, a record or a payment. This is also how you measure accuracy: the reviewer logs what they corrected.

  • Name the reviewer and confirm they have time allocated
  • Log every correction so accuracy can be reported at the end
  • Define what the system may never do without review
  • Decide who can pause or stop it, and how quickly

Exit criteria and what is out of scope

Write down the conditions that end the pilot early: accuracy below a floor, a data problem that cannot be fixed inside the time box, or the owner losing availability. Write down what is out of scope in the same document, so scope creep has to be argued for rather than drifting in.

  • Three or four early-stop conditions, agreed by the owner
  • A short list of adjacent workflows that are explicitly excluded
  • A single decision meeting booked for the end date
  • The three possible outcomes written down: scale, fix, or stop

Next Steps

AI Proof of Concept

How we run a fixed-scope pilot and what it costs.

See the service

AI Use Case Prioritisation Tool

Not sure which workflow to pilot first? Score the candidates.

Prioritise use cases

AI Project Cost Estimator

Estimate the full build cost if the pilot succeeds.

Estimate the cost

Frequently Asked Questions

How long should a first AI pilot run?

Long enough to cover a normal cycle of the workflow and short enough that the owner stays engaged. For most business workflows that means four to eight weeks of live use after setup, with a fixed end date agreed before the start. A pilot that runs longer than a quarter without a decision has usually become an unofficial production system, which is the outcome this checklist is designed to prevent.

What if we cannot measure a baseline?

Then the workflow is not ready for a pilot, and the useful first step is to measure it for two weeks. That can be as simple as a shared spreadsheet where the people doing the work log the count and the time each day. If the workflow genuinely cannot be measured, you will not be able to tell whether the pilot helped, and the decision at the end will be made on impressions.

Who should own the pilot?

The person who owns the workflow today, not the IT team and not the vendor. The owner needs authority to change how the work is done, time to review output and answer questions, and a stake in the result. A pilot owned by someone who does not do the work tends to solve a problem the workers do not have.

Should the pilot use real data?

Yes, a real sample, with the same handling rules you would apply in production. Pilots run on clean synthetic data pass and then fail on the messy real thing. If the data includes personal information, decide before the start how it is stored, who can see it and where it is processed, and apply the same rules the production system would need.

What is a sensible budget cap?

A number the owner can approve without a business case, spent on a fixed scope with a fixed end date. The cap should include internal time, not just an external quote, because the reviewer and the owner will spend real hours on it. Our fixed-price audit is $3,000 plus GST, and a pilot is typically scoped from that audit, so the total is known before work starts.

What happens at the end of the pilot?

One meeting, booked at the start, with three possible outcomes. Scale: the threshold was met, and the pilot becomes a project with a budget and a support owner. Fix: it was close and the gap is understood, so a second short cycle is agreed with a new end date. Stop: it did not work, the reason is written down, and the learning is kept. All three are good results from a pilot. The bad result is no decision.

Items You Could Not Tick?

Send us the workflow, the metric you have in mind and the items that are still open. We will tell you whether it is ready to pilot and what would make it ready.