Contributing expert: Shawn Torkelson, Chief Marketing & Strategy Officer

Tom Stallings, Chief Revenue Officer

 

Companies aren’t failing at AI adoption. Instead, they’re failing to define what AI is supposed to deliver. For the past few years, the dominant success metric in enterprise technology teams has been activity: tokens burned, code shipped, and climbing adoption rates. The assumption is that more AI usage will eventually translate into value. It isn’t, and the evidence is accumulating:

  • Uber’s COO acknowledged that he could find no link between the company’s soaring token spend and any improvement in consumer-facing products.

  • Microsoft canceled licenses for AI coding tools just six months after opening up access, citing cost.

  • A Meta employee ran an internal leaderboard that tracked individual token consumption. The leader burned 281 billion tokens in a single month, at an estimated cost of $1.4 million.

In addition to these high-profile examples, Goldman Sachs research published in early 2026 found no meaningful relationship between AI adoption and economy-wide productivity, citing a 2026 estimate of $700 billion invested in AI that will contribute essentially zero to US GDP growth. Median gains of roughly 30% occurred only where companies were actively measuring the productivity impacts of specific tasks. A recent Forrester Best Practice Report found that 40% of enterprises that launched substantial AI efforts generated no ROI because the outputs never translate into business decisions anyone acts on.

 

The cost of using the wrong AI scorecard

The problem isn’t the pace of adoption; it's the questions organizations are asking. Instead of asking “What value does this create?”, they ask “How much AI are we using?”. Usage is not value. And the gap between the two doesn’t close on its own: it takes a framework built to define what value actually means.

This is the distinction that Coherent Solutions’ Digital Value Creation (DVC) framework is built around. Digital transformations focus on technology change, while DVC focuses first on business impact — using technology to improve earnings before interest, taxes, depreciation, and amortization (EBITDA), scalability, investment readiness, or operational agility, with every initiative tied to measurable outcomes.

DVC starts with an organization’s aspirational outcomes and architects the digital solutions that get them there. The companies pulling back on token spend aren’t retreating from AI. They’re discovering, at a high cost, the use-versus-value gap that DVC was designed to address.

 

The gap between capability and value

DVC uses a specific chain to keep outcome visibility alive across the lifecycle:

Why most AI initiatives dont scale info

Each stage is a conversion step. The chain only pays off at the end, and most organizations break down long before they get there.

The most common fracture point occurs during the experience step between capability and behavior. The capability is there — systems are technically sound, and AI is generating code, summaries, recommendations, and reports. Nothing in the business is actually changing; there are no meaningful improvements in decision speed, customer experience, or cost. The work is happening, but the value isn’t landing.

This is the digital value gap, and it has a recognizable signature. Teams feel busier, not more productive, and activity is up, but outcomes are flat. Additionally, senior engineers are more saturated than before AI arrived, absorbing coordination and review load that the delivery system was never redesigned to handle. All of these concerns roll up to leadership, who sees dashboards full of adoption metrics and feels uneasy about what isn’t on them: the outcome data that would show whether any of it is working.

Why most AI initiatives dont scale info 1

MIT’s 2025 GenAI Divide study found that 95% of enterprise generative AI pilots delivered little to no measurable financial impact. Moreover, companies allotted more than half of their budgets to sales and marketing, while the clearest ROI was in back-office automation, the areas organizations weren’t prioritizing or measuring. The value was there, but they were using the wrong scorecard.

With stats such as these, it’s clear that the value gap doesn’t close by adding more AI. It closes by changing how the delivery system is organized, and that starts with knowing the outcomes you want to achieve.

 

Change the delivery system, change the outcome

For organizations trying to close the AI value gap, they usually focus on the tools. The real answer, however, is structural. Delivery systems need to evolve to support AI’s increased speed and output. Without these changes, common points of failure develop and compound within the system, stalling growth and cost reduction.

Signal decay

In most systems, business intent only travels forward, with reinterpretation at every handoff — through planning, backlog updates, prompts, and generation. By the time work is delivered, the final product is far from the original goal. The further AI-generated output travels from the original business problem; the less it resembles a solution to it.

Validation debt

AI produces output faster than classic delivery systems can review it. Work accumulates in a state where effort has been made, but no one has confirmed whether it solves the right problem or moves the outcome it was meant to move. The faster AI generates output, the deeper this debt grows until review becomes a bottleneck (and frustration) that erases the speed advantage entirely.

Optimization collision

Engineering, security, and quality teams each apply AI within their own lanes. While this accelerates productivity at the team level, it often creates friction at the system level. Faster code generation doesn’t help if review queues, deployment pipelines, and governance checks aren’t redesigned to keep pace.

While these three failure modes occur at different points of the delivery system, they don’t stay separate for long. They compound each other, widening the gap between activity and value. When the Boston Consulting Group analyzed hundreds of AI case studies, it found that 70% of AI’s value comes from people and processes, 20% from technology and data, and just 10% from algorithms. When it comes to successful AI adoption, the tools are the smallest part of the equation. The delivery system and user behavior are where value is found or lost.

 

Why the standard fixes don’t work

When addressing AI-induced system failures, most organizations want to add more to the system. The “more” could include governance, metrics, reviewers, or tools. And while the response is understandable, it’s almost universally counterproductive.

More governance adds controls, but without feedback loops that connect those controls to outcomes, the bottleneck doesn’t shrink; it’s just better supervised. Meanwhile, more metrics focus on tracking token consumption, code volume, and adoption rates. While these metrics quantify activity, they don’t drive change.

For many organizations, adding more reviewers seems like the most intuitive fix. If AI is producing more output than the system normally handles, simply increasing the number of eyes on the review queue should solve the issue, correct? The solution isn’t so simple. Additional reviewers might be counterproductive, as they further complicate the existing coordination problems, accelerating decision-making or uneven AI application.

Each of these common fixes treats the symptoms. None of them treat the structural inefficiencies. Traditional delivery systems are organized around activity, not outcomes, and they were never redesigned to absorb AI’s speed while prioritizing value creation.

 

The structural fix: DVC and the Continuous Delivery Loop

DVC and the Continuous Delivery Loop (CDL) are the two-part answer to the value gap problem in digital initiatives. DVC helps organizations define what outcomes matter — growth, scalability, investment readiness, operational agility — and organize digital initiatives around those outcomes before a line of code is written. CDL restructures the delivery system to determine how those outcomes get consistently delivered without being lost or reinterpreted across handoffs.

DVC operates through three composable capability pillars that work as one connected framework:

Why most AI initiatives dont scale info 2

Underneath these three interdependent layers, CDL is the operating framework that supports each pillar’s outcomes and success.

Where traditional delivery systems organize work as linear phases, CDL reimagines it as a series of compounding loops within four outcome-focused activity centers:

  • Problem Identification: define the business outcome and intent before a line of code is written.

  • Validation & Exploration: test assumptions and prototype fast, so teams commit to the right build before investing in the wrong one.

  • Design & Engineering: build and ship with governance embedded in the workflow, not bolted on at the end.

  • Observation & Scale: measure real outcomes in production and feed what's learned back into the next loop.

Why most AI initiatives dont scale info 3

Business goals are defined before a line of code is written, anchored to the outcomes DVC identified, so the original intent doesn’t decay across handoffs. Feedback moves continuously rather than accumulating in each phase. Meanwhile, governance runs inside the delivery loop, not as a detached layer or an afterthought applied at the end of the process.

 

DVC and CDL in action

A greenfield program initially scoped for 16 engineers was reassessed with CDL-enabled delivery in mind and revised to 4.5 roles, delivered faster and at lower cost. Engineers became orchestrators. Roles blended architecture, product thinking, and agent management. The economics of delivery changed fundamentally because the operating model changed, not the tools.

Planet Fitness shows what the platform layer produces in practice: a 15% increase in dev velocity that translated into 150% user growth in seven months and increased visitor-to-member conversion. While at the data layer, better-targeted integration drove a 75x reduction in API calls and a projected $4.4M annual ad revenue opportunity for a major delivery platform.

 

Baseline metrics: the foundation for DVC and CDL

Before applying DVC and CDL, it’s essential to gather baseline system metrics. Most organizations begin adopting AI without establishing how their delivery systems functioned before. While this allows them to adopt AI faster, it also makes accurate value measurement nearly impossible. Without a starting point, it’s difficult to precisely measure value and improvement. Instead, organizations can only calculate rough efficiency gains that are often forgotten by the next planning cycle.

So where should organizations start? When gathering baseline metrics, four layers need to be measured: delivery economics, experience impact, operational outcomes, and enterprise value indicators. Each layer connects to the next, and none of them work unless the organization can answer a crucial question: “Compared to what?” For example, in the Planet Fitness case study referenced above, Coherent was only able to track the increase in development velocity and user growth because it had baseline metrics to compare with.

Most organizations skip this and try to attribute impact after the fact. In McKinsey’s 2025 State of AI report, they surveyed approximately 2,000 organizations and found that while nearly 90% are regularly using AI, most have not embedded it deeply enough to realize material enterprise-level benefits. The transition from pilots to scaled impact remains a work in progress not because the tools aren’t available, but because the organizational infrastructure to support them isn’t in place.

DVC begins every engagement by defining the metrics that will determine whether value was created. Unfortunately, most leadership teams never formally ask questions that make value determination and metrics definition possible.

Here are the common DVC questions most organizations skip:

  • What specific business KPIs does this initiative move?

  • What behavioral changes are required for that to happen?

  • Do we have baseline metrics to measure results against?

  • What is the cost of delay if we don’t act now?

These are the types of questions that help companies think beyond AI capability to evaluate business intent. And they’re exactly what gets skipped when companies are only concerned with tracking activity metrics.

 

The leadership question underneath it all

When executive teams define desired outcomes before a program launches, every AI initiative can be evaluated against those outcomes from the start rather than audited against them at the end. When outcomes are set prior to launch, the delivery system, tooling, and operating model will follow that initial alignment. Without it, even the best framework won’t be as effective as teams are essentially chasing a vague moving target.

Forward-thinking companies aren’t retreating from AI. Instead, they’re recalibrating around value. The recalibration is useful only if the delivery system is rebuilt to prioritize outcomes rather than activity.

That’s where DVC is most effective. And CDL is the undergirding framework that helps close the digital value gap.

AI pilots are easy.

Scaled impact is an entirely different problem.