Telecommunications organizations are more and more trying to AI to assist groups navigate extremely specialised domains, however generic fashions typically lack the industry-specific information wanted to know telecom networks, requirements, and operations. To handle that hole, AT&T created their Open Telco (OTel) fashions, the following era of telecom-focused AI designed to convey deeper telecommunications experience into AI programs. Constructing OTel2.0 required greater than coaching a big language mannequin, it mirrored a broader challenge many organizations face: how you can construct domain-specific AI programs at scale whereas balancing value, efficiency, and operational complexity. Price administration shortly grew to become a key consideration. To proceed advancing telecom-focused AI, AT&T wanted a platform able to supporting OTel2.0 growth at a completely new scale.
The place groups beforehand needed to personal and handle deployments, infrastructure, and the related operational overhead, Foundry Managed Compute supplied a extra streamlined option to entry devoted graphics processing unit (GPU) capability. This transformation requires greater than highly effective fashions; it requires the power to scale with out compromising value, flexibility, or efficiency.
Utilizing Microsoft Foundry Managed Compute, AT&T was capable of experiment throughout a number of open fashions, optimize workloads throughout completely different GPU architectures, and course of large volumes of telecom knowledge all inside a unified platform. The consequence was an AI growth atmosphere able to supporting trillions of tokens whereas giving groups the flexibleness to iterate, optimize, and innovate quicker.
Mannequin alternative meets infrastructure flexibility
Constructing OTel2.0 required flexibility throughout each fashions and infrastructure. Fairly than standardizing on a single mannequin, AT&T adopted a multi open-model technique. Open fashions had been central to AT&T’s strategy as a result of they supplied the flexibleness to work with accredited telecom knowledge, tailor the workflow for domain-specific mannequin growth, and help large-scale experimentation with better management over value and deployment technique. Via Microsoft Foundry, the staff deployed a number of fashions from the Hugging Face assortment, together with Phi-4, OSS-120B, and Gemma-4, to help completely different phases of growth, from artificial knowledge era and knowledge preparation to reasoning-intensive workloads and broader mannequin growth efforts. Phi-4 performed a major function on this course of, processing greater than 700 billion tokens a month as a part of the broader knowledge preparation and coaching workflow for OTel2.0.
Each firm on the earth must construct its personal AI, and that’s solely potential with open fashions and open supply. AT&T is championing this imaginative and prescient, constructing on open fashions like Phi-4 and Gemma, and giving OTel again to the neighborhood as a telecom AI basis others can construct upon. Microsoft Foundry makes this sensible at scale, bringing the newest open fashions from the Hugging Face assortment along with AMD and NVIDIA GPUs in a single place, so groups can choose the best mannequin and the best {hardware}, then deploy in hours as a substitute of weeks.
—Jeff Boudier, Vice President of Product, Hugging Face
Growing OTel2.0 additionally required infrastructure able to working at telecom scale. AT&T used roughly 530 GPUs by means of Microsoft Foundry Managed Compute spanning a number of GPU architectures together with 430 AMD Intuition™ MI300X GPUs. This heterogenous strategy gave AT&T extra flexibility in how fashions had been deployed and optimized as necessities advanced.
This flexibility illustrates a broader pattern throughout AI growth. Organizations more and more want platforms that enable them to decide on the best mannequin for the job, optimize for value and efficiency, and scale workloads with out rebuilding operational environments. Microsoft Foundry brings mannequin alternative, infrastructure flexibility, governance, and operational scale collectively in a unified platform that helps these necessities.
Past flexibility and price, deployment pace is a important issue for a lot of AI initiatives. As workloads develop and new fashions are evaluated, the power to entry GPU capability shortly allows groups to maneuver from experimentation to execution quicker with out prolonged provisioning cycles. With Foundry Managed Compute, AT&T may deploy and scale fashions in days relatively than ready weeks for infrastructure to develop into out there, serving to speed up growth timelines and preserve momentum throughout OTel2.0 growth.
Optimizing value with out limiting innovation
As AI workloads develop, economics develop into as vital as mannequin efficiency. For AT&T, one of many main aims was to decrease AI mannequin consumption prices whereas persevering with to drive significant enterprise worth by means of AI-powered innovation. By utilizing open fashions on Microsoft Foundry Managed Compute, AT&T was capable of help large-scale knowledge preparation and mannequin growth utilizing a unique financial mannequin constructed round devoted GPU infrastructure and open-model flexibility.
The affect grew to become clear at scale. In help of OTel2.0, AT&T processed roughly 1T tokens, consisting of uncooked paperwork from GSMA supplemented by artificial knowledge generated. Producing the info utilizing open-source fashions like Phi-4, served by Microsoft’s Foundry Managed Compute, saved tens of thousands and thousands of {dollars} versus utilizing frontier fashions. This allowed groups to put money into larger-scale experimentation and growth whereas sustaining a deal with enterprise worth and operational effectivity.
When you find yourself processing a whole lot of billions of tokens, infrastructure turns into a part of the issue you remedy. Foundry Managed Compute gave us entry to GPU capability at scale so our groups may deal with advancing OTel2.0 as a substitute of managing infrastructure.
—Mark Austin, Vice President, Information Science and AI at AT&T
At this scale, infrastructure is now not merely a deployment consideration. It turns into a strategic element of AI growth.
Accelerating the following wave of production-scale AI
OTel 2.0 demonstrates how organizations can mix open fashions, scalable infrastructure, and area experience to construct production-ready AI programs. By matching completely different fashions to completely different workloads and optimizing infrastructure for value and efficiency, AT&T was capable of course of trillions of tokens whereas sustaining operational effectivity.
As organizations transfer from AI experimentation to manufacturing deployment, they more and more want the flexibleness to decide on the best fashions, optimize infrastructure, and scale effectively. Microsoft Foundry and Foundry Managed Compute assist help that transition by bringing these capabilities collectively in a unified platform.
Be taught extra
Discover session subjects from AMD’s Advancing AI:

