Building the Agentic Service Cloud

Building the Agentic Service Cloud

*Blog by Jitin Bhandari, EVP, NeoCloud Platform and Forward Deployment Engineering, Mavenir

The Seven Capabilities That Turn AI Infrastructure into a Trusted Platform 

Where Paper 1 left off

Paper 1 argued that agentic workloads are larger, less predictable and harder to govern than chat, and that the durable opening is not more GPU capacity but the trusted operating and commercial layer through which agentic work is executed. 

This paper asks the next question: what must actually be built? 

The answer is not a better model. Gartner predicts that by 2029 at least 70% of organizations running production agentic AI in infrastructure and operations will suffer a material service, security or cost incident linked in part to insufficient runtime controls. As Gartner puts it, written policies cannot physically stop an agent from making a destructive error.[1]

I call this an Agentic Service Cloud: a focused platform that integrates open infrastructure and model ecosystems while owning the control points customers actually buy — governance, tenancy, evidence and economics. 

Why “full stack” should not mean “build everything” 

The lower layers of the stack are standardizing quickly and in the open: 

Layer What is happening in the open ecosystem Implication
Serving and scheduling llm-d, backed by Red Hat, Google, IBM, NVIDIA, CoreWeave, AMD and others, entered the CNCF Sandbox in March 2026, adding cache-aware routing and prefill/decode disaggregation on top of vLLM[2] Adopt; do not rebuild serving kernels or schedulers
Tool connectivity Anthropic donated the Model Context Protocol to the Linux Foundation’s Agentic AI Foundation in December 2025, with OpenAI, Google, Microsoft and AWS among members[3] Adopt the standard; own the policy around it
Agent-to-agent Google’s A2A protocol now sits under the same neutral governance body[3] Interoperate rather than invent
Models Open-weight and low-cost models lead token volume on major routing platforms Federate catalogs; route by policy and cost

Nutanix’s April 2026 positioning for NeoClouds — production agentic applications, multitenancy, secure self-service, predictable token costs and sovereign deployment — reflects the same shift: value is moving to the layer that makes these components consumable as governed services.[4]

Looking at this from the lens of a CSP or NeoCloud provider, the design principle that follows is selective ownership: 

  • Adopt from open source: serving engines, schedulers, orchestration, agent frameworks, tool protocols, evaluation and guardrail detectors. 
  • Take from partners: facilities, power, accelerators, frontier-model access, payments, tax and general ledger. 
  • Own: tenant hierarchy, policy, evidence, useful-output economics and the commercial relationship.

The Agentic Service Cloud adopts standardizing components and owns the control points that define trust and economics. 

Capability 1: Heterogeneous infrastructure abstraction 

Few CSPs or emerging NeoClouds will own all the capacity they sell. Their estates will combine owned, leased and partner-provided accelerators across NVIDIA, AMD and inference-specific silicon, and across central, regional and edge sites. 

Decode — the token-by-token generation phase of inference — is memory-bandwidth bound, which is exactly where alternative accelerator architectures compete. A platform that assumes a single accelerator family makes an implicit bet on the least settled part of the stack. The platform must therefore: 

  • Present owned, leased and partner capacity as one schedulable pool, with the source recorded for cost and sovereignty. 
  • Distinguish announced, contracted, energized, accepted and revenue-producing capacity; conflating them is a common source of commercial error. 
  • Place workloads by latency, cost, jurisdiction and accelerator fit, not availability alone. 

Capability 2: Inference as a governed service 

Inference is where enterprise data, third-party models and regulatory obligations meet in real time. Fleet-level engineering matters: cache-aware routing and disaggregated serving have shown substantial latency and throughput gains over naive load balancing in reference deployments, and more than 70% of agentic tokens on OpenRouter come from cached prompts. Routing and caching are commercial levers, not just engineering choices.[5]

  • Multi-model interfaces spanning frontier, open-weight and operator-private models. 
  • Policy-based routing deciding per request on cost, latency, quality and jurisdiction. 
  • Useful output as the unit of measure: a token that meets its quality and latency target, not merely a token produced. 

Capability 3: Agent runtime and tool governance 

A long-running agent behaves like a distributed system with credentials. Identity is the first gap. The Cloud Security Alliance reports that non-human identities outnumber human users by about 45 to 1 on average, fewer than one in four organizations have formal policies for creating or removing AI identities, and more than 16% do not track new AI identities at all.[6]

  • A distinct identity for every agent, with delegated, time-bound authority. 
  • Tool authorization at the protocol layer, so every MCP or API invocation is mediated, logged and metered. 
  • Session and memory controls defining what an agent may remember and for how long. 
  • Bounded autonomy: confidence thresholds, human approval for consequential actions, verification and rollback. 

Capability 4: Enterprise knowledge and context 

Agents are only as reliable as the context they act on. As Paper 1 noted, operators name cross-domain integration, data quality and data silos as their top autonomy barriers — data problems, not model problems.[7]

Flat vector retrieval often struggles with questions that depend on hierarchy, precedence and cross-references. One 2026 study on U.S. federal regulations reported a 70% accuracy improvement from knowledge-graph-enhanced retrieval over vector-only RAG; results will vary by domain, but structure and ontology matter.[8]

  • Retrieval, knowledge graphs and domain ontologies encoding organizational context. 
  • Lineage from source data to answer, and provenance for every model, adapter and prompt version. 
  • Protection of tacit enterprise and national knowledge, so customers retain ownership of intelligence built on their data. 

Capability 5: Operator-grade multitenancy 

An operator that becomes an AI provider is simultaneously a customer, a landlord and a channel to enterprises, public bodies and resellers. That requires a nested hierarchy — operator, reseller, enterprise, department, application, agent — in which every level has its own identity, policy, quota, meter and audit boundary, and policy defined once at the top propagates downward. 

Isolation profile Typical use
Shared service with logical isolation Non-sensitive commercial AI
Namespace or virtual-cluster isolation Enterprise departments and applications
Dedicated nodes or accelerators Regulated data where buyers require hardware separation
Dedicated cluster or sovereign enclave Public sector, critical infrastructure
Physically isolated or air-gapped Classified workloads

Most cloud platforms model accounts and projects. Few model an operator reselling governed AI to many downstream tenants on one estate. 

Capability 6: Sovereignty as evidence 

Residency answers where. Sovereignty also answers who can compel access, who can administer the system, what leaves the boundary, and whether the customer can prove it. A practical model spans control domains including data, models, keys, operations, telemetry, support, software updates, legal entity, supply chain, continuity and exit, and independent assurance. 

Regulation is turning this into a record-keeping obligation. Article 12 of the EU AI Act requires high-risk AI systems to support automatic event logging over their lifetime; following the 2026 Digital Omnibus, obligations for stand-alone high-risk systems apply from December 2027.[9]

Capability 7: Agentic economics 

Because agentic consumption is larger and spikier, the commercial layer becomes a control system, not a billing afterthought. 

  • Pre-use authorization. Reserve budget before an agent starts, re-authorize as it runs, and stop it when the budget is exhausted. Telecom online charging has controlled real-time consumption this way for decades; the same semantics now apply to tokens. Many AI billing tools were built invoice-first and meter after the fact. 
  • Token, workflow and outcome metering. Tokens remain the base unit, but buyers increasingly want to pay for resolved incidents, completed transactions or SLA-backed services. 
  • Cost-to-revenue reconciliation. Supplier cost from owned, partner and leased capacity reconciled to revenue by tenant, model and jurisdiction. 
  • Multi-party settlement. Revenue shared, disputed and assured across models, agents, APIs and infrastructure providers. 

The role of Forward-Deployed Engineering 

A platform does not become a production outcome on its own. Indeed data reported by Business Insider shows forward-deployed engineer postings grew about 729% year over year to April 2026 — a signal that deployment, not models, is now the scarce capability. For CSPs and NeoClouds, FDE should:[10]

  • Convert the platform into production outcomes — the first governed workflow, first sovereign tenant, first reconciled invoice. 
  • Build reusable patterns, not permanent customization — each engagement should leave a blueprint the next customer can adopt with less effort. 
  • Feed production learning back into the product — field gaps become the roadmap. 

The measure of a healthy practice is declining customization per customer. If every deployment is bespoke, the provider is running a consultancy, not a platform. 

Note: Frameworks, thresholds and phase timings in this series are the author’s recommendations, not industry benchmarks. Vendor modeling is labeled as such.

Five questions to test a platform strategy 

  1. Which of the seven capabilities do we own today, and which are we implicitly outsourcing? 
  1. Can we give every agent an identity and stop it at runtime? 
  1. Can we host a regulated sub-tenant on the same estate as our own workloads without a second build? 
  1. Can we produce sovereignty evidence on demand, or only in response to an audit? 
  1. Can we authorize spend before an agent runs, and reconcile cost to revenue afterwards? 

What comes next 

Having the platform is not enough. A provider must still decide what to sell first, to whom, at what price, on whose capacity and through which partners. Paper 3 — From Metal to Agentic Revenue turns the platform into a business. 

Which of the seven capabilities is hardest for your organization today? Message me for the one-page capability checklist to benchmark what you should own, adopt or source from partners. 

References

  1. Gartner, Agentic AI Fails Where Governance Stops (Sept 2026) 
  2. CNCF / llm-d project announcements (2026); KubeCon EU 2026 inference stack reporting 
  3. Linux Foundation Agentic AI Foundation announcements; MCP 2026 roadmap 
  4. Nutanix press release, April 7, 2026 
  5. OpenRouter token usage data (2026) 
  6. Cloud Security Alliance: State of Non-Human Identity and AI Security (Jan 2026); The NHI Governance Vacuum (May 2026) 
  7. TM Forum Autonomous Networks survey (2026) 
  8. Chakraborty & Guha, Knowledge Graph RAG, arXiv:2604.14220 (2026) 
  9. EU AI Act, Article 12; Regulation (EU) 2026/1744 (Digital Omnibus) — subject to legal review 
  10. Indeed data via Business Insider (May 2026) 

Dark Mode