On-Premises SLMs
Fine-tuned Small Language Models built from operator data run on operator GPU infrastructure, eliminating per-seat and per-token cloud licensing at scale.
AI is generating tokens at consumer scale. Every inference call, every enterprise prompt, every automated agent task consumes tokens. Operators sit at the exact convergence point of network, compute, and billing.
Mavenir’s Integrated AI Platform transforms that position into a revenue stream. Operators can meter, charge, and monetize AI token consumption on the phone bill, in the same way mobile data plans are billed today, while keeping subscriber data and model weights on their own infrastructure.
The platform runs on operator-controlled infrastructure as its primary deployment model. Prompts, model weights, and subscriber data remain within the operator’s network boundary under normal operations. Where specific use cases require frontier model capability, a policy-governed Model Router provides encrypted connectivity to external APIs, metered through the same charging infrastructure.
Fine-tuned Small Language Models built from operator data run on operator GPU infrastructure, eliminating per-seat and per-token cloud licensing at scale.
General-purpose inference runs on operator GPU infrastructure using open-source foundation models for standard workloads. Cost-efficient and fully controlled.
For tasks requiring advanced reasoning or multimodal capability, the Model Router provides policy-governed, metered access to external frontier models. All usage is billed through MDE.
Each component addresses a specific challenge operators face when deploying AI at carrier scale, spanning economics, routing, billing, security, and sovereignty.
Routes inference requests to the optimal model based on policy, cost, latency, and task complexity across three model tiers: on-premises SLMs, open-source foundation models, and frontier cloud models. Routing decisions are governed by operator-defined policies per plan tier, subscriber segment, and workload type. Frontier model access is metered through the same token charging infrastructure as on-premises inference, providing unified cost visibility and billing regardless of where the model runs.
Reduces token consumption before LLM calls through context pruning, log compression, cache alignment, and metadata elimination. Operators deploying AI services at consumer scale require predictable GPU economics. The Token Optimizer converts variable per-token cloud cost into a manageable on-premises capacity model.
Builds fine-tuned Small Language Models from operator data, including requirements agents, customer service models, network operations assistants, and domain-specific reasoning models. SLMs run entirely on-premises on operator GPU infrastructure. The SLM Builder integrates with Red Hat AI’s MLOps and model registry pipelines.
Mavenir’s Digital Enablement (MDE) provides the token metering, charging, and billing integration layer for operator AI monetization. MDE counts input and output tokens with billing-grade accuracy, enforces plan quotas and overage rules, generates per-session and per-department itemised records, and integrates with operator BSS and mediation systems. Hard caps, threshold alerts, and anomaly detection protect both operators and subscribers from runaway AI spend.
Closed-loop service assurance monitors AI service health with full SLA awareness, distinguishing business-critical AI inference from background workloads and prioritising resources accordingly. Automated fault detection, predictive remediation, and SLA management ensure operators can commit contractual service levels for their operator-managed AI offerings.
Intent-based policy management spans model behavior, network QoS, and compute allocation from a single control plane. Operators define high-level intent, such as priority model routing with guaranteed latency for premium plan subscribers, and the platform translates that intent into coordinated policy across the AI gateway, model router, network slice, and charging system. Policy updates propagate in real time without manual reconfiguration.
The platform runs on operator-controlled infrastructure as its default and primary deployment model. Prompts, model weights, and subscriber data remain within the operator’s network boundary under normal operations. Operators configure which workloads may reach external models, which subscriber tiers are eligible, and what cost thresholds apply. Compliance with data sovereignty, GDPR, and national AI governance frameworks is maintained through configurable data residency controls.
The platform provides Identity, Authentication, Trust, and Authorization services establishing a Zero Trust foundation for AI applications and agents. Verifiable identities, strong authentication, delegated authority, and policy-based access control ensure that every model, agent, tool, and data source operates within explicitly defined permissions. Built-in guardrails are enforced at the model gateway, with continuous monitoring of agent behavior and threat detection tuned to adversarial prompt patterns and model extraction attempts.
Token-based plans give operators a familiar, proven commercial model. Subscribers and enterprises buy AI capacity the same way they buy data, with hard caps, threshold alerts, and anomaly detection protecting both sides from uncontrolled spend.
Entry-level AI access bundled with existing mobile plans. MDE handles metering from day one, with no new billing infrastructure required.
Per-department AI budgets with itemised usage records. Enterprises gain cost visibility; operators gain stickier B2B relationships.
Committed capacity plans with contractual SLAs for large enterprise customers, backed by closed-loop service assurance.
MDE provides the token metering, charging, and BSS integration layer that makes operator AI monetization commercially viable. Hard caps, threshold alerts, and anomaly detection protect both operators and subscribers from uncontrolled AI spend. Operators can bill AI consumption on the phone bill using the same proven infrastructure that already handles mobile data.
Operator infrastructure is the default, primary deployment environment. Under normal operations, all prompts, model weights, and subscriber data remain inside the operator’s network boundary.
Every model, agent, tool, and data source operates within explicitly defined permissions. The AI Security layer adds protection against threats specific to AI environments at carrier scale.
The platform is designed to deliver results operators can report, spanning commercial, operational, and compliance dimensions.
Mavenir’s Integrated AI Platform is built in collaboration with Red Hat, combining Mavenir’s operator monetization, network expertise, and telco-grade SLA management with Red Hat OpenShift AI’s enterprise ML infrastructure and model registry pipelines.
The SLM Builder integrates directly with Red Hat AI’s MLOps pipelines. OpenShift AI provides the container-native runtime that lets operators deploy, manage, and scale AI models with the same operational discipline applied to their core network functions.
The result is a validated, jointly supported platform engineered for the compliance, availability, and operational requirements of tier-one operators.
For more information, visit mavenir.com/mavscale