Akamai has announced an NVIDIA AI Grid-based inference cloud that routes AI requests across edge, regional and core infrastructure. The company says the orchestration layer is designed to weigh latency, GPU availability, throughput and cost when placing distributed inference workloads.
Akamai has introduced AI Grid intelligent orchestration, an NVIDIA AI Grid-based offering intended to route AI inference workloads across the company’s edge, regional and core computing infrastructure.
According to Akamai’s announcement, the service is designed to use its distributed network, including 4,400 edge locations, alongside larger regional and core facilities. The company said the system is powered by thousands of NVIDIA Blackwell GPUs and dynamically directs requests according to requirements such as latency, capacity, throughput and cost.
The offering focuses on inference: the stage at which a trained model processes an input to produce an answer, prediction or other output. For interactive AI applications, network delay can materially affect the user experience. Akamai’s approach is intended to place requests closer to users when response time is the priority, while sending other work to regional or core resources when capacity or economics are more important.
Akamai said its orchestration layer can make workload-placement decisions among its infrastructure tiers rather than relying on a single central location for inference. In principle, that could give organizations building geographically distributed AI services more flexibility in handling changing demand and constrained GPU resources.
The company describes the platform as an inference cloud that balances latency, cost and throughput. Its announcement does not provide detailed technical information on the routing algorithms, supported model families, developer interfaces or service-level commitments.
IDC, in an analyst report published by Akamai, characterized the deployment as a global implementation of NVIDIA AI Grid. The report said the orchestration control plane dynamically directs inference requests based on latency, cost and GPU availability.
Computer Weekly likewise reported that Akamai’s implementation spans edge, regional and core infrastructure. The publication framed the launch as an effort to treat inference capacity as a distributed resource that can be routed across a global network, rather than as compute fixed in one type of data center.
The announcement reflects a broader infrastructure challenge for AI providers and enterprises: applications may need low response times near users, but accelerated compute is costly and not always available in every location. A distributed model aims to combine nearby processing for latency-sensitive requests with access to deeper pools of GPU capacity elsewhere.
For Akamai, the launch extends the role of its existing network beyond content and application delivery. The company is positioning that footprint as a foundation for serving AI inference workloads whose placement can change with user location, performance needs and available compute.
The practical value of the platform will depend on how consistently its orchestration layer can match requests to suitable GPU resources while maintaining predictable performance. Akamai has not detailed pricing, broad availability timing or which AI models customers can deploy through the service in the cited materials.
Akamai’s AI Grid initiative illustrates how network operators are seeking to make AI inference more geographically flexible. By connecting edge, regional and core resources under a common routing layer, the company aims to offer different latency and cost options for production AI services.
According to Akamai’s announcement, the service is designed to use its distributed network, including 4,400 edge locations, alongside larger regional and core facilities.
The company said the system is powered by thousands of NVIDIA Blackwell GPUs and dynamically directs requests according to requirements such as latency, capacity, throughput and cost.
The offering focuses on inference: the stage at which a trained model processes an input to produce an answer, prediction or other output.
Continue reading