A CEO once asked our team for something simple: a mapping of where AI could responsibly be applied inside the business. We went to the people running the processes and asked a basic question. How long does this activity take?
For one activity, the answers ranged from 6 hours to 120 hours. A factor of twenty. If the true number is 6, it’s a minor task. If it’s 120, it’s three full working weeks. Nobody was lying. Nobody was hiding anything. The range existed because the activity had never been formally measured. There was no before.
The lack of a documented “before” is the reason most companies cannot tell you what AI actually returned. Not because the technology failed. Because nobody can prove a change without a starting point.
This isn’t one company’s problem. More than half of the chief executives who signed the AI checks say the technology delivered no revenue benefit and no cost benefit. That’s from PwC’s 29th Global CEO Survey, published in January 2026, which asked over 4,000 chief executives across 95 countries what AI had actually done for their revenue and costs in the prior 12 months: 56% reported neither. Only 12% reported both a revenue gain and a cost reduction (PwC, 2026). These are not the opinions of skeptics. These are the people who approved the budgets.
Other studies land in the same place through different methods. MIT’s NANDA Initiative found that about 95% of enterprise generative AI deployments show no measurable impact on profit and loss (MIT NANDA, 2025). IBM found that only 25% of AI initiatives delivered the expected return, and only 16% scaled enterprise-wide (IBM, 2025). In 2024, BCG found that 74% of companies had yet to show tangible value from AI (BCG, 2024). McKinsey’s latest State of AI report found that 6% of organizations are AI high performers, attributing at least 5% of EBIT impact to AI use and having seen significant value from AI. Meanwhile, only 37% of organizations reported that AI has positively contributed to their EBIT, a share that remains unchanged from the previous year (McKinsey, 2026). Different samples, different years, different questions. The conclusion is the same.
The critical distinction is buried in that data. Most companies are not only failing to get value. They are failing to prove it. Those are two different problems, and the second one is more fixable.
Why Companies Buy Before They Measure
The instinctive explanation for poor AI returns points to talent gaps, data quality, or technical complexity. The data points somewhere else entirely. In a late 2023 Gartner survey across the United States, Germany, and the United Kingdom, the top barrier to AI adoption, cited by almost half of respondents, was the difficulty in estimating and demonstrating the value of AI projects. That ranked above talent shortages, above technical difficulty, above data issues, and above trust concerns (Gartner, 2023).
Rita Sallam, Distinguished VP Analyst at Gartner, put it plainly: “Executives are impatient to see the returns on GenAI investments, yet organizations are struggling to prove and realize the value.” (Gartner, 2024). The operative word is “prove.” Proof requires a before. Most organizations have never created one.
What makes this pattern self-reinforcing is an admission embedded in IBM’s 2025 study: 64% of CEOs acknowledged that the risk of falling behind drives investment in some technologies before they have a clear understanding of the value those technologies bring. Nearly two-thirds are buying first and understanding later (IBM, 2025). The purchase decision is driven by competitive anxiety, not by a documented hypothesis about which specific activity will get faster, and by how much.
Deloitte’s 2026 State of AI in the Enterprise found that 37% of companies report using AI only at the surface level (Deloitte, 2026). The license exists. Deep use does not follow automatically from its existence.
This kind of gap is common. In a Prime BPM survey of business process management professionals, 49% said their organization had started creating process maps without a defined process architecture, and 48% said their organization lacked a defined business process improvement methodology (Prime BPM). Most companies aren’t behind. They’re normal, and normal isn’t ready to measure AI.
The Competitive Gap Is Methodological, Not Technical
The companies that cannot demonstrate AI returns are not necessarily getting no returns. That distinction matters operationally. If the return exists but cannot be pointed to, the problem is not the technology. The problem is the absence of a baseline: a documented number that records how long a given activity took before the tools arrived.
Without a baseline, a measurement-first competitor can do things a license-first organization cannot. It can report the number of hours removed per activity, the cost per transaction, and the payback period. It can rank use cases by the size of the baseline they attack. It can negotiate with vendors from a position of knowledge rather than hope. It can stop a failing pilot after eight weeks with a number rather than a feeling. Same tools, same people, different starting conditions.
A license-first organization, by contrast, can report adoption rates, satisfaction scores, and anecdotes. All real. None of them is a saving. It cannot prioritize among use cases because it lacks a common unit of comparison. It cannot kill a failing pilot because it cannot tell that the pilot is failing.
What to Watch
Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. The reasons cited were poor data quality, inadequate risk controls, escalating costs, and unclear business value (Gartner, 2024). The last item on that list is the one a baseline directly addresses.
Watch for the moment the CFO asks for the return on the AI investment. In a license-first organization, that question triggers a search for a baseline that does not exist. The team then produces activity metrics: adoption percentages, usage counts, and employee satisfaction scores. These are reported as if they were financial outcomes. They are not. The CFO either accepts the substitution or escalates. Either response signals an organization that hasn’t yet solved the measurement problem.
The gap first becomes visible to practitioners during process mapping. When a team asks employees how long a given activity takes, the answers reveal whether the organization has ever formally measured the work it is now trying to accelerate. A wide range of estimates is not a personnel issue. It is a structural signal.
The Sequence That Fails, and the One That Works
The sequence that fails has five steps, and every step is individually reasonable. The problem is not any single step. The problem is the order.
- Buy licenses. Often driven by competition, a department pilot is usually launched.
- Hope. Roll out the pilot and hope that employees’ productivity improves. There is training involved, perhaps a town hall or champions program to encourage usage.
- ROI is asked. A senior leader asks for the ROI three to six months later.
- Search for a baseline. The team looks for a baseline and discovers there is none. No one had measured how long activities took before AI.
- Report activity metrics. Activity metrics such as “Employees feel more productive” or “Adoption is at 60%” replace savings.
The sequence that works also has five steps, and the first three cost no license fee.
- Select. Select three to five high-volume, repetitive processes.
- Measure. Using a single consistent method, determine the hours per activity, the frequency, and who performs the activity.
- Baseline. Get a documented number with a source, date, and owner.
- Introduce AI. Introduce artificial intelligence into mapped processes responsibly, with a clear hypothesis about which activity it will shorten.
- Measure again. Remeasure using the same method and unit, and report the difference as the saving.
This is not a new framework. It is the same logic applied to any capital investment. What makes AI different in practice is that the tools are inexpensive enough and the vendor claims loudly enough that many organizations believe the standard of proof doesn’t apply. It does.
The baseline is an asset. The license is a cost. One of those two items depreciates without the other.
Why Employees Can’t Just Tell You the Number
The economics of the measurement gap become clear once quantified. A case from a recent engagement illustrates the problem precisely. A CEO asked for process mapping before purchasing new tools. The team selected specific processes, went to the employees who performed them, and asked a simple question: How much time does this specific activity take within this process? For one activity, the answers ranged from 6 hours to 120 hours. That’s a factor of 20.
If the true number is 6 hours, the activity is a minor task. If it’s 120 hours, it represents three full working weeks at 40 hours per week. Try to build a business case within that range. What does AI save? 2 hours? 40 hours? Both are technically correct answers given the range, which means neither is usable.
The research explains why employees can’t close that range from memory. Buehler, Griffin, and Ross, writing in the Journal of Personality and Social Psychology in 1994, asked students to predict how long it would take to finish their honors thesis. The average best estimate was 33.9 days. The actual average was 55.5 days. Only 29.7% finished within their own best estimate. When asked for a worst-case estimate, 48.7% finished even by that pessimistic date (Buehler et al., 1994). That’s the planning fallacy. People estimate based on an idealized inside view and systematically ignore variability and interruptions. The width of the range is not dishonesty. It is a structural feature of how humans estimate work they’ve never formally timed.
The problem compounds in modern work environments. Microsoft’s 2025 Work Trend Index, surveying 31,000 knowledge workers across 31 countries, found employees are interrupted roughly every two minutes during core hours, and that 60% of meetings are unscheduled (Microsoft, 2025). An activity is rarely done in one sitting. It’s started, interrupted, resumed, and finished days later. The elapsed time and the actual touch time become two different, equally true numbers. Neither gets measured.
None of this means employees are bad at estimating in general. A 1998 study in the BLS Monthly Labor Review found that self-reported total weekly hours are reasonably reliable in aggregate, correlating strongly with diary-based measures (Jacobs, 1998). The problem isn’t honesty. It’s estimating one specific activity without measurement.
From AI Adoption to AI Accountability
The category is moving from “Do you have AI tools?” to “Can you prove what those tools returned?” The first question is nearly settled across large enterprises. The second is the one replacing it.
That shift changes what operating leaders need to produce. Adoption rates satisfied early sponsors because they signaled movement. CFOs and boards are now asking for the financial line item: what cost came down, what capacity was freed, and which transaction became cheaper? Those questions cannot be answered without a pre-implementation measurement. Deloitte’s finding that 37% of companies remain at a surface-level use of AI (Deloitte, 2026) suggests that a large population hasn’t yet faced the accountability question. For those organizations, the question is coming. The companies that build measurement infrastructure now will have an answer. The ones that don’t will produce the same activity reports already failing to satisfy executives in the more advanced cohort.
The Homework
The one question worth asking before signing the next license agreement is simple: What is the documented baseline for the process this tool is supposed to improve? Not, “What productivity gain does the vendor claim?” Not, “What are peer companies reporting?” The specific, dated, sourced measurement of how long the target activity takes today, in this organization, performed by these people.
If that number doesn’t exist, the next step isn’t to purchase. It’s to measure. Select three to five high-volume, repetitive processes. Time them using a consistent method, not recalled estimates. Record the source and the date. That document costs no license fee, and it’s the only thing that makes the subsequent AI investment measurable, prioritizable, and defensible.
The organizations that get this right won’t necessarily have better AI tools than their competitors. They’ll have something more durable: a before. And without a before, there is no after, which means there is no return anyone can point to.
Works Cited:
- https://www.pwc.com/gx/en/issues/c-suite-insights/ceo-survey.html
- https://www.artificialintelligence-news.com/wp-content/uploads/2025/08/ai_report_2025.pdf
- https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles
- https://www.bcg.com/press/24october2024-ai-adoption-in-2024-74-of-companies-struggle-to-achieve-and-scale-value
- https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- https://www.gartner.com/en/newsroom/press-releases/2024-05-07-gartner-survey-finds-generative-ai-is-now-the-most-frequently-deployed-ai-solution-in-organizations
- https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
- https://www.deloitte.com/content/dam/assets-zone3/us/en/docs/services/consulting/2026/state-of-ai-2026.pdf
- https://web.mit.edu/curhan/www/docs/Articles/biases/67_J_Personality_and_Social_Psychology_366,_1994.pdf
- https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born
- https://www.bls.gov/opub/mlr/1998/12/art3full.pdf
- https://www.primebpm.com/global-bpm-trends
