This weblog put up is the primary of a four-part collection known as The Economics of Agent Optimization which shares the methods, capabilities, and proof factors that can assist you optimize agent prices and run AI as a managed funding system on Microsoft Foundry.
The AI dialog in most enterprises has moved from the whiteboard to the price range assessment. Two years in the past, the query was whether or not AI might work. The query leaders are asking now could be sharper and fewer comfy: is it paying for itself?
For the groups now in manufacturing—together with greater than 100,000 organizations constructing on Microsoft Foundry that query has change into pressing. Tokens have change into the brand new unit of know-how spend, and monetary self-discipline (not mannequin alternative) is what decides whether or not a promising pilot ever scales. The cash is already shifting in: in a Microsoft-commissioned IDC research of greater than 4,000 enterprise leaders, 71% mentioned they plan to extend AI budgets, funded from IT and non-IT sources alike. The budgets are rising. The query is whether or not the self-discipline grows with them.
71% of enterprise leaders plan to extend their AI budgets
2025 IDC survey
The groups pulling forward didn’t go searching for a less expensive mannequin. They stopped working AI as a string of one-off pilots and began working it as a managed funding system: each request sized to its job, each agent improved because it runs, and each greenback bounded and accounted for. That shift, from shopping for intelligence to managing it, is the entire sport. This collection is about how the system works and why Microsoft Foundry is constructed to run it.
Perceive your AI prices and spending
Earlier than you may handle AI spend, you must perceive what creates it. Price shouldn’t be decided solely by the mannequin you select. Additionally it is formed by the appliance or agent constructed round that mannequin.
Each request contains enter tokens, corresponding to system prompts, dialog historical past, software definitions, and retrieved content material, in addition to output tokens generated by the mannequin. As a result of fashions are stateless, the complete context is distributed with each request. Prices can improve over time even when the person asks solely a easy follow-up query.
Brokers introduce one other layer of complexity. As a substitute of following a single path, an agent could consider choices, retry actions, or name a number of instruments earlier than producing a response. A single person request can generate many mannequin calls, making workflow design as essential as mannequin choice.
Enhance AI value visibility throughout groups
AI spend is tough to handle when it seems as a single mixture quantity. Groups want visibility into prices by software, agent, workflow, and mannequin to grasp what’s driving utilization and the place optimization alternatives exist.
With out that stage of attribution, it turns into tough to clarify prices, prioritize enhancements, or measure the affect of optimization efforts.
Management and optimize spend
Visibility alone shouldn’t be sufficient. AI workloads can scale rapidly, and sudden conduct can improve consumption in a brief time frame. Organizations want controls that assist handle spend earlier than prices change into a shock.
Optimization additionally requires greater than deciding on a lower-cost mannequin. Most AI workloads comprise a mixture of requests with totally different necessities. Higher outcomes come from matching requests to the best fashions, lowering pointless context, limiting unneeded software use, and enhancing agent workflows in order that they function extra effectively.
Why Microsoft is the platform for AI FinOps
FinOps started because the self-discipline of bringing monetary accountability to variable cloud spend, a shared working mannequin that places engineering, finance, and product on one set of numbers. FinOps for AI comes all the way down to 4 commitments:
Make AI predictable to fund
Environment friendly by design
Optimized at scale
Confirmed in worth
Microsoft’s reply is a single, first-party strategy to FinOps for AI that spans your entire lifecycle—plan, construct, handle, and measure. Price visibility and management are constructed into the merchandise groups already use: Microsoft Foundry and GitHub the place brokers are constructed and run, Microsoft Price Administration for allocation and chargeback, Azure pricing presents for commitment-based financial savings, and Azure API Administration because the gateway that meters and governs AI visitors. Microsoft Agent 365 extends the identical self-discipline to the tenant—unifying agent value administration throughout Microsoft and third-party platforms with spending insurance policies, price range caps, and departmental chargeback in a single place. Collectively they provide organizations one thing no level software can: complete, best-in-class value administration throughout the entire AI property, from the primary immediate to the board-level ROI quantity.
Foundry is the place that strategy will get particular, as a result of it’s the place brokers are run and optimized. It runs AI as a managed funding system throughout one closed loop: optimize every request at runtime, optimize every agent workflow over time, and govern the spend constantly.
AI value optimization begins with visibility
A managed funding system makes three choices, every at a distinct velocity. You optimize the request in the second it runs. You optimize the agent workflow over days and weeks, as you be taught what works. And also you govern the spend constantly, with limits and budgets that by no means sleep. Foundry is constructed to make all three. Every transfer has its personal set of Foundry capabilities, and the map beneath exhibits how they match collectively.
Agent 365 will prolong governance to the tenant, unifying value administration throughout Microsoft and third-party brokers with spending insurance policies, price range caps, and departmental chargeback.
You’ll be able to watch the runtime levers work dwell in our new Microsoft Mechanics episode on token economics.
The 4 questions AI leaders must be asking
If you happen to take one factor from this put up, take these 4 questions into your subsequent AI or price range assessment. Every has a concrete reply in Foundry. If you happen to can’t reply one at present, that’s the place to begin.
Do we all know what we’re paying for?Spend must be seen by mannequin, agent, and workflow, not hidden in a single bill line. Foundry’s metering and traces make it simpler to grasp the place prices originate.
Are we paying the correct amount for every request?Most requests don’t want a frontier mannequin. Mannequin router, deployment and pricing choices, caching, fine-tuning, and Foundry IQ assist match every request to the aptitude it wants.
Are our brokers working effectively?Agent prices ought to enhance over time as workflows change into more practical. Agent optimizer and reminiscence in Foundry Agent Service and Toolboxes in Foundry assist cut back pointless token utilization and enhance execution high quality.
Do our limits maintain when utilization spikes?Utilization that expands quickly wants controls that maintain. In the present day, many groups put Azure API Administration in entrance of their AI endpoints to implement token charge limits and quotas on the AI Gateway layer. Native budgets and enforcement inside Foundry, plus tenant-wide controls via Agent 365, are the place we’re headed subsequent.
The primary query is about understanding AI spend. The following three are the areas this collection explores in additional element: matching requests to the best fashions, enhancing agent effectivity, and making use of governance controls to handle value at scale.
Get began
This collection will proceed over the approaching weeks, going one stage deeper on every subsequent transfer: methods to optimize the request at runtime, methods to construct brokers that use tokens effectively, and methods to govern the spend as you scale. Every put up pairs the considering with the Foundry capabilities that make it actual.
You don’t have to attend to begin. The capabilities behind this framework are dwell in Microsoft Foundry at present:
Comply with alongside because the collection unfolds and deliver the 4 inquiries to your subsequent assessment.

