Beyond GPU-as-a-Service
*Blog by Jitin Bhandari, EVP, NeoCloud Platform and Forward Deployment Engineering, Mavenir
Why the Agentic Economy Opens a New Cloud Opportunity for CSPs and NeoClouds
From generation to execution
The first phase of generative AI followed a simple pattern: a user sent a prompt, a model returned an answer, and the user decided what to do next.
Agentic AI changes that pattern. The user sets an objective. The system plans the steps, retrieves information, selects models and tools, works across applications, checks results and continues until the task is complete or it reaches a decision boundary.
| A chatbot generates an answer. An agentic system executes a workflow and delivers an outcome. |
Consumer launches have made the shift visible. Meta launched Muse on September 8, 2026: a personal agent running in a dedicated cloud virtual machine that can open a browser, complete forms, send email and make purchases on a user’s behalf. It joins a fast-growing category of persistent agents, including Grok Bot, Town, Instinct and Google’s Gemini Spark.[1]
Consumption data shows the same trend. On OpenRouter, agentic token usage grew about 14x in six months between February and August 2026, while human-driven usage grew 2.8x.[2]
Enterprises will create the larger economic impact, but adoption is still ahead of execution:
| Indicator | Data point |
|---|---|
| Enterprise applications with task-specific agents | 40% by end-2026, up from under 5% in 2025 (Gartner forecast)[3] |
| Leaders saying they are adopting agentic AI | Three-quarters; only a small minority in meaningful production (Forrester, 2026)[4] |
| Most-funded AI use cases meeting revenue expectations | About one in four (ISG, 2025)[4] |
| Agentic projects expected to be canceled by end-2027 | More than 40%, due to cost, unclear value or weak risk controls (Gartner, June 2025)[3] |
The gap between ambition and production is where the opportunity sits.
| The Agentic Economy will need more than models and megawatts. It will need trusted operators. |
Why agentic workloads change cloud economics
Generative AI is measured by the cost and quality of an answer. Agentic AI must be measured by the cost, reliability and business value of a completed task.
A single objective may cause an agent to make many model calls, retrieve enterprise data, invoke tools and APIs, coordinate with other agents, maintain memory, request approval and replan after a failure. Every step consumes tokens, and context accumulates with every loop.
Anthropic’s engineering team has quantified the multiplier from its own production system: a single agent uses roughly 4x the tokens of a chat interaction, and a multi-agent system roughly 15x. Observed ranges vary widely by workload; academic work on coding agents also shows the same task can vary by up to 30x in token use between runs.[5].
Three consequences follow for infrastructure providers.
1. Demand is structurally larger. Deloitte estimates inference reaches about two-thirds of all AI compute in 2026, up from roughly half in 2025.[7]
2. Demand is less predictable. When the same task can vary 30x in token cost, budgets built for pilots break in production. Consumption must be authorized and bounded, not only billed after the fact.[6]
3. Cost control becomes a commercial capability. Where tokens are processed, on which model and at what price is now a commercial decision, not only a technical one.
| Owning an AI factory is not the same as operating an AI services business. |
An AI factory converts energy and compute into tokens. A NeoCloud business turns those tokens into services that can be provisioned, governed, measured and sold. The same accelerator can produce materially different revenue depending on whether it is sold by the hour or as governed, metered services – a point Paper 3 quantifies.
The scaled NeoClouds are already acting on this. CoreWeave’s FY2025 annual report shows $5.1 billion of revenue alongside a $1.2 billion net loss and $60.7 billion of remaining performance obligations, and the company has acquired platform and developer-tooling businesses. Nebius has stated a strategy to extend beyond bare-metal compute into software and services. Capacity alone is not being treated as a defensible position.[8]

As lower layers standardize or commoditize, differentiation concentrates in the control plane between capacity and outcomes.
Why the opportunity is not limited to hyperscalers
Hyperscalers start with major advantages: scale, the broadest service catalogs, strong developer ecosystems and established enterprise procurement relationships. Analysys Mason estimates they held about 76% of GPU-as-a-service revenue in 2024 — but forecasts that share falling to about 63% by 2030.[9]
Enterprise agentic systems will run across multiple clouds, private environments, edge locations and regulated jurisdictions. Some tasks will need frontier models; many will run on open or specialized models close to enterprise data, where economics are better.
At inference time, models work with live customer records, contracts and proprietary knowledge. Sovereignty must therefore cover more than storage location: where models run, who administers them, who holds the keys, where telemetry goes and what evidence proves compliance.
The telecom industry has history to learn from. McKinsey notes that global mobile data traffic grew about 60% a year between 2010 and 2023 while telecom revenue grew about 1% a year — value accrued largely to the platforms above the network.[10]
This creates four credible openings:
- Jurisdictional trust. Providers that can produce evidence of control over data, models and administration will be stronger than those offering residency alone.
- Distributed execution. NVIDIA reports 18 telcos across five continents launched sovereign AI factories within 18 months, with future growth expected in distributed edge inference.[11]
- Multi-provider neutrality. A platform that routes across models, capacity sources and jurisdictions preserves choice while applying one set of governance rules.
- Industry context. Valuable enterprise agents need operational context, industry data models and trusted tools, not just general-purpose reasoning.
CSPs and NeoClouds have relevant assets – but no automatic right to win
Owning connectivity does not create a NeoCloud. Owning data centers does not create an AI platform. Deploying individual agents does not create an agentic business.
Agents also need reliable operational truth. TM Forum autonomous-network survey data shows about 75% of operators still at Level 1–2 autonomy and only about 4% at Level 4 in any single domain. The barriers operators cite are data problems, not model problems: cross-domain integration (60%), data quality (54%) and data silos (54%).[12]
| For CSPs | For emerging NeoClouds | |
| Structural advantage | National trust, enterprise and public-sector channels, edge footprint, identity and real-time charging | Speed, AI-native engineering, accelerator expertise, freedom from legacy |
| Typical gap | AI platform engineering, developer experience, model and inference operations | Enterprise governance, sovereignty evidence, vertical depth, channel reach |
| Strategic choice | Buyer, channel, infrastructure provider, sovereign AI operator, platform or marketplace | Capacity utility, managed platform, sovereign provider, vertical cloud or orchestrator |
| Biggest risk | Building capacity without a monetization layer | Competing only on GPU price as capacity commoditizes |
The underdeveloped layer: orchestration plus economics
A single enterprise task may call a frontier model, a local model, enterprise knowledge services, communications or payment APIs, another agent, and network and accelerator resources. Each raises a commercial question:
- Who authorizes the workflow, and which budget applies?
- How is each provider’s contribution measured?
- How is cost reconciled with revenue?
- Does the customer pay per token, workflow, resolution or outcome?
- How is revenue settled among participants?
Security raises the stakes. Agents act with real credentials, and CrowdStrike reports AI-enabled attacks rose 89% year over year while average breakout time fell to 29 minutes. Identity, authorization and evidence are commercial requirements, not optional features.[13]
TM Forum argues that today’s agent marketplaces are largely hyperscaler-governed, leaving CSPs as consumers rather than economic owners, and that the larger opportunity is a CSP- or consortium-owned plane combining agent lifecycle, policy, execution and economic settlement.[12]
| The strategic control point is the environment that decides which agent or model runs, under what authority, in which jurisdiction, for which tenant, at what cost, with what evidence and under which commercial agreement. A provider does not need to own every model or accelerator to hold it. |
| Author’s perspective Mavenir builds cloud-native software that operators run in production networks worldwide (data, voice & messaging large scale systems), real-time charging and service assurance. Those systems already solve, for sovereign complexity, many of the problems agentic AI now raises: real-time authorization before consumption, nested operator and reseller hierarchies, multi-party settlement and carrier-grade evidence. My team’s NeoCloud and forward-deployed engineering work applies that operating discipline to AI. That is the lens behind this series. |
Note: Frameworks, thresholds and phase timings in this series are the author’s recommendations, not industry benchmarks. Vendor modeling is labeled as such.
Five questions for CSP and NeoCloud leaders
- Where do we have a structural right to participate? Jurisdiction, enterprise access, operational data, distributed infrastructure, charging capability or industry expertise?
- Which layer do we intend to own? Infrastructure, inference, agent platform, commercial control plane, marketplace or vertical application?
- What is our first repeatable, measurable outcome?
- Can we prove trust? Where execution occurred, which model acted, what data was used and who authorized it.
- Can we connect consumption to economics? When the same task can vary 30x in token cost, this is a margin question.
What comes next
Paper 2 — Building the Agentic Service Cloud: the seven capabilities that turn infrastructure into a trusted platform.
Paper 3 — From Metal to Agentic Revenue: positioning, supply, catalog, unit economics, forward-deployed engineering and go-to-market sequencing.
Market opening → Platform design → Commercial execution
Which of the five questions is hardest for your organization to answer today? I’d welcome your view in the comments — and the conversation with CSP, NeoCloud and technology leaders working through it.
References
- Meta Muse launch coverage: TechCrunch, CNET, SiliconANGLE (Sept 2026)
- OpenRouter agentic vs. human token usage data (Feb–Aug 2026)
- Gartner: task-specific agents in enterprise applications forecast; Gartner press release, June 25, 2025 (agentic project cancellations)
- Forrester, The State of Agentic AI, 2026; ISG, State of Enterprise AI Adoption, 2025
- Anthropic Engineering, “How we built our multi-agent research system” (June 2025)
- Stanford / University of Michigan study on agentic coding token consumption
- Deloitte TMT Predictions 2026
- CoreWeave Form 10-K FY2025; Nebius Group Form 20-F FY2025
- Analysys Mason, GPUaaS worldwide forecast 2025–2030 (Oct 2025)
- McKinsey, AI infrastructure: a new growth avenue for telecom operators (2025)
- NVIDIA, Telcos Across Five Continents Are Building Sovereign AI Infrastructure (May 2025)
- TM Forum Autonomous Networks survey (2026); TM Forum, The Agentic AI Marketplace Opportunity (June 2026)
- CrowdStrike 2026 Global Threat Report