Most AI programmes are not blocked by a lack of technical options. They are blocked by an inability to show, in commercial terms, which problem deserves attention first. That matters because funding decisions made without a baseline turn plausible demonstrations into expensive indecision. The decision should change: measure the work before choosing the system.
The pattern we keep seeing is simple. Several initiatives appear credible, each has an executive sponsor, and none can establish a defensible comparison with the current process. The organization postpones the choice, calls the delay prudence, and absorbs the cost of leaving the work unchanged.
This is not an argument against experimentation. It is an argument about sequence. A pilot can demonstrate that a model produces a useful answer. It cannot, by itself, show that the answer improves a workflow enough to justify implementation, controls, training and ongoing operating cost.
The difficult part is usually not calculating the price of an application. It is describing the work accurately. How many cases arrive? How often are they reworked? Which decisions require a senior employee? Where does waiting create customer loss, revenue leakage or regulatory exposure? Without those observations, an AI business case quietly substitutes enthusiasm for evidence.
Consider a commercial insurer deciding whether to automate claims intake. A prototype can extract information from submitted documents and route straightforward cases. That is a compelling demonstration. The investment decision depends on what happens around it: the time adjusters spend correcting incomplete submissions, the proportion of claims that need escalation, the delay caused by handoffs, and the cost of errors that reach a customer.
Suppose the insurer discovers that extraction is not the main constraint. Most delay comes from missing information and a review queue owned by another team. Improving document intake may make one step faster while leaving the customer-facing outcome unchanged. A less impressive intervention, such as changing the intake form or routing rules, could have greater value and lower risk.
That is why measurement is also a prioritization mechanism. It separates work that is expensive because it is repetitive from work that is expensive because it is uncertain, poorly designed or dependent on another decision. AI is useful in the first category only when the surrounding process can absorb the change. In the second, better process design may be the more valuable intervention.
There is a reasonable objection: early measurement can slow teams down, while learning often requires building something. True. A baseline should not become a six-month analysis exercise. It should be proportionate to the decision. For a contained experiment, a sample of cases, a time estimate and an agreed success measure may be enough. For a customer-facing or regulated workflow, the baseline must also include error costs, escalation paths and human accountability.

External evidence points in the same direction, though the estimates are not uniform. Gartner predicted that at least 30 percent of generative AI projects would be abandoned after proof of concept because of poor data quality, weak risk controls, rising costs and unclear business value. The precise rate matters less than the failure pattern: a demonstration is not evidence of an operating capability.
The larger consequence is organizational. When teams measure the work first, AI funding becomes a choice about capacity, risk and service quality, not a contest between compelling prototypes. The question is no longer which initiative looks most advanced. It is which change can be judged against the work the organization is already paying for.
