Why Enterprise AI ROI Is An Architecture Problem
Ragy Thomas is Chairman, Co-CEO and Co-Founder of UnifyApps, and Founder and Chairman at Sprinklr. Author - βThe Enterprise Brain.β
gettyβCIOs are continually asking me the same two questions: What is AI actually costing us, and is it delivering measurable ROI?
The numbers explain the anxiety. According to Gartner, β72% of CIOs reported that their organizations are breaking even or are losing money on their AI investments.β Gartner also forecasts that βover 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs.β
AI isnβt getting expensive because intelligence itself costs more. It is getting expensive because most businesses bolt AI onto fragmented architecture. Every new agent rebuilds context, integrations, controls and verification around itself, while workflows become hard-wired to individual models. That same fragmentation makes ROI harder to prove because costs, decisions and outcomes remain scattered across systems.
Fix the architecture and you change both sides of the ROI equation: The marginal cost of deploying AI falls, while the economic contribution of each workflow becomes easier to measure. Below are four architectural moves that can improve the unit economics of enterprise AI.
Most AI agents become expensive because they have to fetch context repeatedly from multiple systems just to answer one question.
One multinational beauty companyβs team saw this with their first attempt at agentic automation. The models worked and demoed well. But the agents kept breaking the moment they hit production, because each of the six point solutions theyβd stitched together kept their own knowledge and governance in a silo. Every agent started cold, forcing the team to reconnect, reconcile and rebuild that context for every single new project.
The answer is to centralize that context instead of rebuilding it agent by agent. Create a shared context layer that unifies enterprise knowledge, governance rules and permitted actions, so every AI workflow can draw from the same reconciled foundation rather than reconnecting to the business from scratch.
Context is already prepared and shared, which cuts redundant token usage, reduces API traffic and makes responses both faster and cheaper.
Shared context lowers the marginal cost of each new deployment because knowledge, permissions and business logic are built once and reused. At the same time, it also makes ROI easier to attribute by linking workflow cost, actions and outcomes in one place.
When every step in a workflow runs on the same model, simple tasks end up carrying the cost of complex reasoning. One financial-services enterprise studied by researchers from the University of Hong Kong and Stellaris AI found that inference costs exceeded $200,000 per month even though more than 70% of queries were routine enough for smaller models.
When you treat intelligence as a swappable input, models can change without rebuilding the surrounding workflow.
Businesses should optimize high-value, judgment-intensive tasks for performance and optimize high-frequency, low-value tasks for cost. Where a smaller model or deterministic logic does the job, use it.
This way, inference spend becomes intentional, controlled and tied to the real complexity of the work. The goal is predictable economics and model agility.
Rather than routing every task through a single, monolithic βsuperagentβ that must hold the full context of a problem in memory and reason over it end-to-end, a swarm of specialized agents breaks the work into narrower, well-defined tasks. Each agent operates with a smaller, purpose-built context window instead of the sprawling prompt a generalist model would need to handle every possible scenario.
Because these agents are only invoked when relevant, and only pass forward the specific output another agent needs, the system avoids the redundant token overhead of reprocessing full context at every step.
The result is a meaningful drop in token cost per task, along with faster response times, since specialized agents typically need fewer reasoning tokens to arrive at a correct answer within their narrow domain.
Notably, the efficiency only holds if the swarm is governed as one system, with shared permissions, controls and observability across agents. With those guardrails in place, specialized agents can reduce token overhead without creating a new layer of operational risk.
Governance gets expensive when itβs added one application at a time. Every new agent brings another round of access controls, approvals, audit requirements, escalation rules and human verification.
Mckinsey found retrofitting governance after AI development can create costly rework and deployment delays. Gartner estimates that βeffective governance technologies could reduce regulatory expenses by 20%.ββ
Consider a production AI claims workflow. It may need role-based access to policyholder data, audit trails, limits on autonomous actions and clear escalation rules. Build 10 agents independently, and much of that control framework must be designed, tested, documented and monitored 10 times.
Building governance into the platform flips that equation. Define the controls once, then apply them as code across every workflow, agent and integration. Compliance overhead grows more slowly because controls are reused instead of rebuilt.
The same architecture turns the AI black box into a glass box. Each workflow leaves a clear record of the data used, actions taken, human intervention required, cost incurred and business outcome delivered. That visibility makes ROI easier to measure because leaders can see not just what AI costs, but what economic value each workflow creates or protects.
Enterprise AI will not become economically sustainable through lower token prices alone. The better path is to pick one use case with a clear business outcome, take it into production, identify the capabilities that make it work and build those capabilities to be reused elsewhere.
Capabilities are the architectural primitive. Use cases are where those capabilities meet pain points.
Over time, the economics improve. Shared context reduces repeated integration work. Model flexibility keeps inference costs optimized. Platform-level governance keeps control costs flatter. Reusable capabilities mean each new deployment starts with more already built.
The leadership question is, βWhich use case should we take into production first, and what should we build underneath it so the next one is cheaper, faster and easier to measure?ββ
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


