Friday · October 9, 2026

 ·  Daily  ·  Newsletters

Resonance network: Quantum InsiderSpace Insider

SoftBank Corp. Announces ‘Infrinia AI Cloud OS,’ a Software Stack for AI Data Centers

Insider Brief

  • SoftBank Corp. says its Infrinia Team has developed Infrinia AI Cloud OS, a software stack aimed at helping AI data center operators rapidly deploy GPU cloud services spanning model training through inference while lowering operational complexity and total cost of ownership.
  • The platform enables Kubernetes as a Service and Inference as a Service in multi-tenant environments, automating infrastructure from hardware configuration to Kubernetes orchestration and offering OpenAI-compatible APIs for large language model inference.
  • SoftBank plans to deploy Infrinia AI Cloud OS within its own GPU cloud services first, with plans to expand to overseas data centers as part of a broader push to position itself as a global provider of next-generation AI infrastructure software.

PRESS RELEASE — SoftBank Corp. (TOKYO:9434, President & CEO: Junichi Miyakawa, “SoftBank”) announced that its Infrinia Team*1, which works on the development of next-generation AI infrastructure architecture and systems, has developed “Infrinia AI Cloud OS,” a software stack*2 designed for AI data centers.

By deploying “Infrinia AI Cloud OS,” AI data center operators can build Kubernetes*3 as a Service (KaaS) in a multi-tenant environment, and Inference as a Service (Inf-aaS) that provides Large Language Model inference capabilities via APIs, as part of their own GPU cloud services. In addition, the software stack is expected to reduce total cost of ownership (TCO) as well as operational burden compared with bespoke solutions or in-house development. This will enable the rapid delivery of GPU cloud services that efficiently and flexibly support the full AI lifecycle—from AI model training to inference.

SoftBank plans to deploy “Infrinia AI Cloud OS” initially within its own GPU cloud services. Furthermore, the Infrinia Team aims to expand deployment to overseas data centers and cloud environments with a view to global adoption.

Background of “Infrinia AI Cloud OS” Development

The demand for GPU-accelerated AI computing is expanding rapidly across the generative AI, autonomous robotics, simulation, drug discovery, and materials development fields. As a result, user needs and usage patterns for AI computing are becoming increasingly diverse and sophisticated, and requirements including the following have emerged:

  • Access to infrastructure that is fully managed by GPU cloud service providers, abstracted GPU bare-metal servers
  • Cost-optimized, highly abstracted inference services without concerning with GPU management
  • Advanced operations in which AI models are trained and optimized on centralized servers and deployed for inference at the edge

Building and operating GPU cloud services that meet these requirements requires highly specialized expertise and involves complex operational tasks, placing a significant burden on GPU cloud service providers.

To address these challenges, the Infrinia Team developed “Infrinia AI Cloud OS,” a software stack that maximizes GPU performance while enabling the easy and rapid deployment and operation of advanced GPU cloud services.

Key Features of “Infrinia AI Cloud OS”

Kubernetes as a Service

  • Reduces the operational burden of managing the physical infrastructure and the Kubernetes software layer by automating the entire stack (from BIOS and RAID settings to the OS, GPU Drivers, networking, Kubernetes Controllers and Storage) on state-of-the-art GPU Platforms such as NVIDIA GB200 NVL72
  • Software-defined dynamic, on-the-fly physical connectivity (NVIDIA NVLink) and memory (Inter-Node Memory Exchange) reconfiguration, as the customers create, update and delete their clusters to suit their AI workload needs
  • Automatic node allocation based on GPU proximity and NVIDIA NVLink domain to reduce latency and maximize GPU-to-GPU bandwidth for highly distributed jobs

Inference as a Service

  • Enables users to deploy inference services simply by selecting Large Language Models, without working with Kubernetes or the underlying infrastructure
  • OpenAI-compatible APIs, enabling drop-in integration with existing AI applications
  • Seamless scaling across multiple nodes in core and edge platforms such as NVIDIA GB200 NVL72 and other platforms

Secure Multi-tenancy and High Operability

  • Tenant isolation through encrypted cluster communications and separation
  • Automation of operational maintenance, including system monitoring and failover
  • API environment for connecting to the AI data center’s portal, customer management systems, and billing systems

These key features allow AI data center operators with customer management systems, as well as enterprises offering GPU cloud services, to add advanced capabilities that enable efficient AI model training and inference while flexibly utilizing GPU resources, to their own GPU service offerings.

Junichi Miyakawa, President & CEO of SoftBank Corp., commented:
“To further deepen the utilization of AI as it evolves toward AI agents and Physical AI, SoftBank is launching a new GPU cloud service and software business to provide the essential capabilities required for the large-scale deployment of AI in society. At the core of this initiative is our in-house developed ‘Infrinia AI Cloud OS,’ a GPU cloud platform software designed for next-generation AI infrastructure that seamlessly connects AI data centers, enterprises, service providers and developers. The advancement of AI infrastructure requires not only physical components such as GPU servers and storage, but also software that integrates these resources and enables them to be delivered flexibly and at scale. Through Infrinia, SoftBank will play a central role in building the cloud foundation for the AI era and delivering sustainable value to society.”

For more information on “Infrinia AI Cloud OS,” please visit the website below:
https://infrinia.ai/

Greg Bock
About the author
Greg Bock

Greg Bock is an award-winning investigative journalist with more than 25 years of experience in print, digital, and broadcast news. His reporting has spanned crime, politics, business and technology, earning multiple Keystone Awards and a Pennsylvania Association of Broadcasters honors. Through the Associated Press and Nexstar Media Group, his coverage has reached audiences across the United States.

Trending today

a purple and green background with intertwined circles
Business & Markets · Enterprise

OpenAI Brings Interactive Intelligent UI to ChatGPT With GPT-6 as Common Sense Media Rates ChatGPT for Teens an Unacceptable Risk

Business & Markets

Cal AI Co-Founder Zach Yadegari Secures $10M for Persona, an AI Assistant With a Wearable Band

Technology & Infrastructure · Data centres & compute

Economist Estimates AI Buildout Would Need $3.55 Trillion in Annual Revenue

Business & Markets

Hadrian Raises $40M to Tackle the AI Hacking Cyber Security Crisis

Technology & Infrastructure · Applications & agents

Microsoft Reveals Nvidia RTX Spark-Powered Surface Laptop Ultra Built to Run AI Agents Locally

The AI economy, every weekday morning

The daily briefing on LinkedIn. Free, one tap to follow.

Exclusives

Exclusive

South Korea’s AI G3 Strategy: Decoded

Scale-ups to Watch

10 Switzerland-Based AI Scale-Ups You Need to Know in 2026

network, blockchain, digital, hand, web, community, artificial, intelligence, steering, interfaces, bokeh, future, digitization, transformation, change, blockchain, blockchain, blockchain, blockchain, blockchain, transformation
Exclusive

Why Crypto Could Be AI’s Payment Layer: BlackRock Sees Stablecoins Connecting Commerce and Compute

AI Predictions
Exclusive

Why AI Predictions Often Get The Technology Right But The Timeline Wrong

Scale-ups to Watch

10 CEE & Baltics-Based AI Scale-Ups You Need to Know in 2026

More in Technology & Infrastructure

Latest from the same section
Technology & Infrastructure · Data centres & compute

Economist Estimates AI Buildout Would Need $3.55 Trillion in Annual Revenue

7 hours ago
Business & Markets · Startups

Veir Raises $110M in Series C Funding to Scale Superconducting Power Systems for AI Data Centers

21 hours ago
a blurry photo of a colorful object
Technology & Infrastructure · Data centres & compute

Google Sends Its First TPU Into Orbit to Test Space-Based AI Compute

4 days ago
Technology & Infrastructure · Data centres & compute

AWS Drops Data Center NDAs and Open Sources a Jev-Style AI Decision Model

4 days ago

The AI economy, every weekday morning

The daily briefing plus the weekly Scale-ups to watch edition. Free, no spam, unsubscribe any time.