Insights
    Program DeliveryAI Strategy

    Planning for a Number You Don't Have Yet

    6 July 2026·5–6 min read
    Planning for a Number You Don't Have Yet

    The question that stops the room

    A steering committee, somewhere in Belgium, reviewing a proposal to automate invoice data extraction. The sponsor has presented the business case, the timeline, the vendor shortlist. Then the CFO asks the obvious question: "What accuracy will it reach?"

    The delivery lead gives the only honest answer available: "Somewhere between 75 and 95 percent. We'll know after three weeks with your actual invoices."

    In most steering committees, that answer sounds like evasion. Funding requests are supposed to come with commitments attached. And so begins a small, familiar tragedy: either the team invents a number to get through the gate, or the project loses out to a competing proposal whose sponsor was more comfortable inventing one.

    Deterministic methods, probabilistic work

    Classic delivery governance grew up around work whose outcome is decided before it is built. When you commission an approval workflow, you write down the rules, someone builds them, and testing confirms the system does what the specification says. Performance is a design choice. If the spec says three-day turnaround, you can hold someone to three days.

    An AI system inverts this. Whether extraction accuracy on your supplier invoices lands at 82 or 96 percent is not a decision anyone gets to make. It is a property of your documents, your suppliers' formatting habits, your scanning quality, and the model, and it can only be measured. Nobody in the room has the number, and no amount of seniority produces it.

    This is the single most consequential planning fact about AI, and most project methodologies have no slot for it: core performance is discovered, not specified. A stage gate that asks "will the system do what we said it would?" has no meaningful answer at gate one, because nobody can yet say what the system will do.

    The two standard workarounds, both bad

    The first workaround is to promise a number anyway. The business case says 90 percent because the template had a field for it and 90 sounded defensible. If reality comes in at 84, the team either quietly redefines how accuracy is measured or reopens the business case mid-delivery. Both erode trust, and the erosion is unfair to everyone involved, because the original number was fiction and everybody half knew it.

    The second workaround is to refuse to commit. This is more honest and works out worse. Open-ended exploration budgets are the natural habitat of the pilot that never graduates: without a defined finish line, the project runs until enthusiasm or budget runs out, whichever comes first.

    Buy the number before you buy the system

    There is a third option, and it requires reframing what the first phase of an AI project is for. Phase one is not a small version of the build. It is a purchase of information.

    The structure is simple. A fixed, modest budget, typically €25,000 to €60,000 depending on data access headaches. A fixed window, usually three to six weeks. And a deliverable that is explicitly a measurement, not software: the accuracy figure on a representative sample of your real data, the cost per processed item, and a breakdown of where and how the system fails.

    That last part matters more than the headline number. "86 percent accurate" is less useful than "86 percent accurate, and nearly all failures are handwritten annotations and one supplier's invoice template". The first is a verdict. The second is a plan.

    Decide the thresholds before you see the result

    Here is the discipline that separates this from an ordinary pilot: the steering committee writes down its decision rules before the measurement exists.

    Something like: below 80 percent, we stop and bank the learning. Between 80 and 92, we deploy with human review on every item and re-measure in production. Above 92, we automate with sampled review. The exact numbers come from the economics of the process, and they will be different for every use case.

    Why decide in advance? Because after six weeks of work, momentum votes for continuing. A team that has just spent a month with the data will always find a reason the disappointing number could improve. Pre-committed thresholds are how the organisation protects itself from its own sunk costs, and they turn the gate review from a debate into a reading.

    A gate that asks a different question

    Funded this way, the programme still has gates, still has tranches, still gives finance the control it needs. What changes is the question each gate asks. Instead of "is the project on schedule?", the gate asks "did the assumption we funded survive contact with the data?"

    And this reframing rescues the most misunderstood outcome in AI delivery: the early kill. When a feasibility phase comes back at 68 percent against an 80 percent threshold and the project stops, that is not a failed project. That is the governance working exactly as designed. You bought certainty for €40,000 instead of discovering the same fact at month nine, which is the most expensive place to learn it. The write-up should read like a successful purchase, because it was one.

    What the committee gets in return

    A steering committee that adopts this model gives up the comfort of confident numbers at gate one. What it gets back is considerable: business cases built on measured ranges instead of invented points, kill decisions that happen at the cheap end of the project, and delivery teams who can afford to tell the truth in the room.

    That last one is worth more than it sounds. Most of the fictions in AI programmes are not told out of malice. They are told because the governance had no way to receive an honest answer. Build the slot for "we don't know yet, and here is exactly what it costs to find out", and you will be surprised how often it gets used.


    Facing a business case that needs a number nobody has yet? Get in touch to talk through how to structure the feasibility phase, or start with our piece on why most pilots never reach production.