{"id":2000,"date":"2026-07-07T15:20:00","date_gmt":"2026-07-07T15:20:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/"},"modified":"2026-07-07T19:59:04","modified_gmt":"2026-07-07T19:59:04","slug":"foundry-managed-compute","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/","title":{"rendered":"Hugging Face Fashions on Foundry Managed Compute"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<div class=\"not-prose\">\n<div class=\"SVELTE_HYDRATER contents\" data-target=\"BlogAuthorsByline\" data-props=\"{&quot;authors&quot;:[{&quot;author&quot;:{&quot;_id&quot;:&quot;61df1bc442b8fffa092441f6&quot;,&quot;avatarUrl&quot;:&quot;\/avatars\/421af1377c858f6137b6ad9089fc624c.svg&quot;,&quot;fullname&quot;:&quot;Manoj Bableshwar&quot;,&quot;name&quot;:&quot;manojsb&quot;,&quot;type&quot;:&quot;user&quot;,&quot;isPro&quot;:false,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;isUserFollowing&quot;:false},&quot;org&quot;:{&quot;_id&quot;:&quot;5e6485f787403103f9f1055e&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/1583646260758-5e64858c87403103f9f1055d.png&quot;,&quot;fullname&quot;:&quot;Microsoft&quot;,&quot;name&quot;:&quot;microsoft&quot;,&quot;type&quot;:&quot;org&quot;,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;plan&quot;:&quot;enterprise&quot;,&quot;followerCount&quot;:20683,&quot;isUserFollowing&quot;:false}},{&quot;author&quot;:{&quot;_id&quot;:&quot;688797c1629e0ef013c2556c&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/no-auth\/Fo0ehdRXafrOg0jqWuCaw.png&quot;,&quot;fullname&quot;:&quot;Osi&quot;,&quot;name&quot;:&quot;ositanachi&quot;,&quot;type&quot;:&quot;user&quot;,&quot;isPro&quot;:false,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;isUserFollowing&quot;:false},&quot;org&quot;:{&quot;_id&quot;:&quot;5e6485f787403103f9f1055e&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/1583646260758-5e64858c87403103f9f1055d.png&quot;,&quot;fullname&quot;:&quot;Microsoft&quot;,&quot;name&quot;:&quot;microsoft&quot;,&quot;type&quot;:&quot;org&quot;,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;plan&quot;:&quot;enterprise&quot;,&quot;followerCount&quot;:20683,&quot;isUserFollowing&quot;:false}}],&quot;translators&quot;:[],&quot;proofreaders&quot;:[],&quot;lang&quot;:&quot;en&quot;}\">\n<div class=\"not-prose\">\n<div class=\"mb-12 flex flex-wrap items-center gap-x-5 gap-y-3.5\">\n<div class=\"flex items-center font-sans leading-tight\"><span class=\"inline-block \"><span class=\"contents\"><img decoding=\"async\" class=\"rounded-full! m-0 mr-2.5 size-9 sm:mr-3 sm:size-12\" alt=\"Manoj Bableshwar's avatar\" src=\"https:\/\/huggingface.co\/avatars\/421af1377c858f6137b6ad9089fc624c.svg\"\/><\/span> <\/span> <\/div>\n<div class=\"flex items-center font-sans leading-tight\"><span class=\"inline-block \"><span class=\"contents\"><img decoding=\"async\" class=\"rounded-full! m-0 mr-2.5 size-9 sm:mr-3 sm:size-12\" alt=\"Osi's avatar\" src=\"https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/no-auth\/Fo0ehdRXafrOg0jqWuCaw.png\"\/><\/span> <\/span> <\/div>\n<\/div><\/div>\n<\/div>\n<\/div>\n<p>At Microsoft Construct 2026, we introduced Foundry Managed Compute and Hugging Face fashions on Foundry \u2014 a curated catalog of open-weight fashions from the Hugging Face ecosystem, refreshed weekly, deployable in a single click on onto Foundry Managed Compute. Weights are pre-staged in Azure, runtimes are constructed and scanned by Microsoft, and each mannequin within the Assortment ships with the identical enterprise safety, governance, observability, and billing that applies to each different mannequin on Foundry.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe Platform: Microsoft Foundry and Managed Compute<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Microsoft Foundry is a platform for constructing and working agentic AI purposes. Foundry begins with the widest mannequin choice on any cloud \u2014 fashions from Microsoft, OpenAI, Anthropic, Meta, Mistral, DeepSeek, Hugging Face, and others, spanning frontier, open-source, and customized weights \u2014 all accessible by way of a single endpoint and a single set of SDKs in Python, C#, JavaScript, and Java.<\/p>\n<p>On prime of these fashions sits the Foundry Agent Service: multi-agent orchestration with built-in reminiscence, information grounding by way of Foundry IQ, and a catalog of connectable instruments by way of agentic protocols, so brokers can work with enterprise knowledge. As soon as brokers are operating, Foundry supplies end-to-end tracing, real-time monitoring, steady evaluations, and a immediate optimizer that improves agent habits primarily based on eval outcomes \u2014 observability and high quality loops which might be a part of the platform.<\/p>\n<p>Alongside that, builders get entry to:<\/p>\n<p>Content material security filters<br \/>\nActivity-adherence guardrails<br \/>\nAn AI Pink Teaming Agent for adversarial testing<br \/>\nUnified RBAC<br \/>\nNon-public networking<br \/>\nAzure Coverage integration straight inside the platform<\/p>\n<p>Alongside pay-per-token (lowest-friction path to get began) and provisioned throughput (predictable, high-performance manufacturing workloads on frontier fashions), Foundry Managed Compute is the third deployment possibility in Foundry: a managed GPU platform-as-a-service for open-source and customized fashions.<\/p>\n<p>You deploy a mannequin occasion described by the issues that matter to your workload \u2014 parameter rely, context size, and whether or not you need to optimize for latency or throughput \u2014 and Foundry handles the GPU topology beneath, whether or not the occasion lands on one accelerator or a number of, so that you assume and plan in mannequin phrases.<\/p>\n<p>Microsoft takes care of the machine: container updates, runtime upgrades, and safety patches occur routinely on the supported runtimes \u2014 vLLM, SGLang, TensorRT-LLM, NIM, TEI, llama.cpp \u2014 with out redeploying your mannequin, whereas mannequin configuration, deployment habits, and routing stick with you.<\/p>\n<p>That consistency carries by way of the developer floor \u2014 pay-per-token, provisioned throughput, and Managed Compute share:<\/p>\n<p>A single endpoint<br \/>\nThe identical SDKs<br \/>\nThe identical authentication<br \/>\nThe identical observability<br \/>\nA single invoice<\/p>\n<p>Open-source fashions combine with Foundry Brokers the identical manner frontier fashions do, so you&#8217;ll be able to combine mannequin sorts in a single agent and not using a separate integration path.<\/p>\n<p>Managed Compute gives:<\/p>\n<p>International deployments \u2014 broadest capability and greatest pricing<br \/>\nKnowledge Zone deployments \u2014 residency and sovereignty<\/p>\n<p>Similar code, similar workflow. Quota is aligned to accelerator households, so a plan constructed on the H100 household immediately carries ahead as new {hardware} generations come on-line.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhy Hugging Face<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Hugging Face is the general public sq. of open AI: 15 million builders, 400,000 organizations, and over 3 million open fashions revealed, with new frontier capabilities \u2014 agentic coding, video segmentation, speech, embeddings \u2014 touchdown weekly. It is the GitHub of open fashions, the place the neighborhood publishes weights, writes mannequin playing cards, compares evaluations, and pulls fashions for experimentation.<\/p>\n<p>Open fashions have closed the hole with proprietary fashions on benchmark after benchmark, and so they unlock issues proprietary endpoints cannot:<\/p>\n<p>State-of-the-art is now open. Main open-weight fashions are aggressive with the highest closed frontier fashions on essentially the most broadly used benchmarks.<br \/>\nDeep customization. Full weights make it potential to fine-tune, distill, quantize, and adapt with LoRA \u2014 tailoring fashions to your area, your knowledge, and your latency and value targets.<br \/>\nYour mannequin, your internet hosting. Weights run in your tenant on infrastructure you management, behind your inference endpoint, together with your identification and community boundaries.<br \/>\nValue shaping. Pay for accelerators by the hour, scale to zero when idle, and right-size GPUs to the particular mannequin \u2014 helpful for regular, high-volume, or latency-sensitive workloads the place per-token pricing is tougher to foretell.<br \/>\nModel management. Pin a particular mannequin model, consider it, deploy it, and transfer ahead or roll again by yourself launch cadence.<\/p>\n<p>The catch has all the time been the operational layer: discovery, license evaluate, safety screening, runtime choice, GPU sizing, picture constructing, CVE patching, and standing the mannequin up behind an enterprise-grade endpoint. Hugging Face, by itself, will not be an enterprise serving platform. Hugging Face fashions on Foundry is that operational layer, run by Microsoft.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tHugging Face Fashions on Foundry<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The Hugging Face Assortment brings a curated subset of fashions straight into the Foundry Mannequin Catalog:<\/p>\n<p>Refreshed weekly \u2014 trending fashions from the Hugging Face ecosystem are added constantly because the neighborhood publishes them.<br \/>\nEach modality \u2014 textual content, imaginative and prescient, audio, and multimodal: LLMs and VLMs for chat and brokers, ASR and speech translation, embeddings, segmentation, picture technology.<br \/>\nSafetensors solely, no untrusted code \u2014 each mannequin within the Assortment is security-screened and ships within the SafeTensors weight format, with no trust_remote_code execution paths except rigorously reviewed.<br \/>\nThe appropriate runtime for the mannequin \u2014 vLLM and SGLang for LLMs, TensorRT-LLM and NIM the place relevant, TEI for embeddings, llama.cpp for CPU \u2014 Foundry picks the engine that matches the mannequin.<\/p>\n<p>Out of your facet, an open-weight mannequin within the Hugging Face Assortment appears and behaves like another mannequin within the Foundry Mannequin Catalog, and each mannequin within the Assortment has been put by way of a multi-stage publishing pipeline earlier than it ever exhibits up there.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe Curation Pipeline<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Hugging Face and Microsoft work collectively to carry the preferred open-weight fashions from the Hugging Face ecosystem to Microsoft Foundry \u2014 production-ready for enterprise environments \u2014 by way of a scientific curation course of:<\/p>\n<p>Determine trending fashions within the Hugging Face ecosystem \u2014 primarily based on neighborhood indicators, accomplice requests, and buyer demand \u2014 and choose candidates for enterprise readiness.<br \/>\nDisplay for compliance and safety \u2014 mannequin licenses are reviewed towards Microsoft&#8217;s enterprise distribution coverage (with license metadata captured and preserved on the catalog mannequin card), and repositories are inspected for trust_remote_code patterns and customized executable code; any mannequin that will require executing third-party Python at load time is both remediated or excluded.<br \/>\nConstruct, scan, and publish runtimes \u2014 Microsoft builds inference container photos on supported runtimes (vLLM, SGLang, TensorRT-LLM, NIM, TEI, llama.cpp), scans them for CVEs, and indicators and publishes them to a Microsoft-managed container registry.<br \/>\nAdd weights to safe Azure storage \u2014 mannequin weights are pulled from Hugging Face as soon as, validated towards the revealed mannequin card, and saved in Microsoft-managed Azure storage within the areas the place the mannequin is served.<br \/>\nValidate and publish to the catalog \u2014 each mannequin + runtime + accelerator mixture is examined for API conformance (chat completions, embeddings, rerank, and so forth.) and efficiency (latency, throughput, time-to-first-token, inter-token decode time), then the validated mannequin \u2014 with its templates, runtime photos, and weights \u2014 is revealed to the Foundry Mannequin Catalog with a one-click deploy path onto Managed Compute.<\/p>\n<p>As a result of weights are pre-staged in Azure storage and runtime photos stay in a Microsoft-managed registry, your deployments will not want outbound community entry to Hugging Face Hub \u2014 you&#8217;ll be able to deploy to manufacturing inside a non-public community.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tMannequin Runtimes<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Hugging Face fashions on Foundry are powered by a flexible assortment of community-built, open-source inference runtimes \u2014 every chosen and tuned for Foundry Managed Compute, and matched to the mannequin architectures it serves greatest. Throughout all runtimes, the systematic curation course of means new variations and patches land on Foundry rapidly, and current mannequin deployments are upgraded routinely \u2014 with out requiring you to redeploy.<\/p>\n<p>vLLM \u2014 the default high-throughput serving engine for open giant language fashions, tuned for manufacturing GPU workloads. As a result of Hugging Face is a direct contributor to vLLM, any mannequin within the Transformers library can run on vLLM out of the field \u2014 so when a brand new mannequin lands on Hugging Face, it may be served on Foundry the identical day, with no ready on a customized integration.<\/p>\n<p>SGLang \u2014 a serving engine for language and multi-modal fashions, with robust assist for structured outputs (JSON, regex, grammar-constrained technology) that agentic and tool-using workloads rely upon. Hugging Face and the SGLang group have constructed a Transformers backend integration for SGLang, so any mannequin within the Transformers library runs on SGLang out of the field \u2014 and reaches Foundry the identical day it lands on Hugging Face.<\/p>\n<p>Textual content Embeddings Inference (TEI) \u2014 the runtime for embedding, reranker, and sequence-classification fashions. Accelerator-specific photos ship with kernels compiled for every GPU and CPU household Foundry helps, preserving the embedding sizzling path lean for RAG and semantic-search workloads.<\/p>\n<p>llama.cpp \u2014 the CPU and small-GPU path for GGUF-quantized fashions. Helpful for cost-optimized deployments, smaller fashions, and CPU-only areas, with the identical OpenAI-compatible API as vLLM and SGLang.<\/p>\n<p>TensorRT-LLM and NIM \u2014 used on NVIDIA {hardware} the place NVIDIA&#8217;s optimized kernels and Triton-based serving ship meaningfully higher latency or throughput for particular mannequin households.<\/p>\n<p>hf-serve \u2014 Hugging Face&#8217;s personal multi-model inference server, used for mannequin architectures exterior the LLM and embedding quick paths (imaginative and prescient, audio, segmentation, and different Transformers-native pipelines) so the Assortment can cowl each modality with a constant serving layer.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tDeploying and Scoring an Open-Weight Mannequin<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The Hugging Face Assortment within the Foundry Mannequin Catalog is the place you begin, and deployment is 5 steps:<\/p>\n<p>Browse the catalog and decide a mannequin \u2014 the deploy wizard additionally surfaces the mannequin id, deployment template id, and acceleratorType you will want in the event you&#8217;re scripting the deploy by way of SDK or REST.<br \/>\nSelect a deployment template \u2014 latency- vs throughput-optimized, accelerator household, context size, quantization.<br \/>\nConfigure occasion rely \u2014 scale throughput by including mannequin cases.<br \/>\nDeploy \u2014 from the portal, CLI, SDK, or REST.<br \/>\nRating by way of the unified Foundry endpoint with the SDK you already use.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tDeployment Templates<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>A deployment template is the unit of alternative in step 2: a named, versioned asset that pins the runtime, the accelerator household and rely, the context size, and the runtime-specific tuning wanted to serve the mannequin properly \u2014 so choosing a template is the one knob you flip for &#8220;how do I need this mannequin to run.&#8221;<\/p>\n<p>qwen3-32b, for instance, ships with 4 templates the deploy wizard exposes facet by facet:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Template<br \/>\nRuntime<br \/>\nAccelerator<br \/>\nContext<\/p>\n<p>qwen\u2013qwen3-32b\u201340k-nvidia-a100<br \/>\nvLLM<br \/>\n1 \u00d7 A100 80 GB<br \/>\n40K<\/p>\n<p>qwen\u2013qwen3-32b\u201340k-nvidia-h100<br \/>\nvLLM<br \/>\n1 \u00d7 H100 80 GB<br \/>\n40K<\/p>\n<p>qwen\u2013qwen3-32b\u2013128k-nvidia-2xa100<br \/>\nvLLM<br \/>\n2 \u00d7 A100 80 GB<br \/>\n128K<\/p>\n<p>qwen\u2013qwen3-32b\u2013128k-nvidia-2xh100<br \/>\nvLLM<br \/>\n2 \u00d7 H100 80 GB<br \/>\n128K<\/p>\n<\/div>\n<p>Every template arrives pre-tuned for the mannequin \u2014 runtime settings, tool-call and reasoning parsers, scoring path, well being probes, request concurrency, and any model-specific context-extension settings are all set by Microsoft, with any trade-offs referred to as out inline within the template description. Once you script the deploy, you reference the template and Foundry handles the remaining.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tDeploy \u2014 Python SDK<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p><span class=\"hljs-keyword\">from<\/span> azure.identification <span class=\"hljs-keyword\">import<\/span> DefaultAzureCredential<br \/>\n<span class=\"hljs-keyword\">from<\/span> azure.mgmt.cognitiveservices <span class=\"hljs-keyword\">import<\/span> CognitiveServicesManagementClient<\/p>\n<p>shopper = CognitiveServicesManagementClient(DefaultAzureCredential(), SUBSCRIPTION_ID)<\/p>\n<p>deployment = shopper.managed_compute_deployments.begin_create_or_update(<br \/>\n    resource_group_name=RESOURCE_GROUP,<br \/>\n    account_name=ACCOUNT_NAME,<br \/>\n    deployment_name=<span class=\"hljs-string\">&#8220;qwen3-32b&#8221;<\/span>,<br \/>\n    useful resource={<br \/>\n        <span class=\"hljs-string\">&#8220;sku&#8221;<\/span>: {<span class=\"hljs-string\">&#8220;title&#8221;<\/span>: <span class=\"hljs-string\">&#8220;GlobalManagedCompute&#8221;<\/span>, <span class=\"hljs-string\">&#8220;capability&#8221;<\/span>: <span class=\"hljs-number\">1<\/span>},<br \/>\n        <span class=\"hljs-string\">&#8220;properties&#8221;<\/span>: {<br \/>\n            <span class=\"hljs-string\">&#8220;mannequin&#8221;<\/span>: <span class=\"hljs-string\">&#8220;azureml:\/\/registries\/azure-huggingface\/fashions\/qwen&#8211;qwen3-32b\/variations\/1&#8221;<\/span>,<br \/>\n            <span class=\"hljs-string\">&#8220;deploymentTemplate&#8221;<\/span>: <span class=\"hljs-string\">&#8220;azureml:\/\/registries\/azure-huggingface\/deploymenttemplates\/qwen&#8211;qwen3-32b&#8211;40k-nvidia-h100\/labels\/newest&#8221;<\/span>,<br \/>\n            <span class=\"hljs-string\">&#8220;acceleratorType&#8221;<\/span>: <span class=\"hljs-string\">&#8220;H100_80GB&#8221;<\/span>,<br \/>\n        },<br \/>\n    },<br \/>\n).consequence()<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tRating \u2014 OpenAI SDK<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>The deployment is reachable by way of the unified Foundry endpoint with the OpenAI SDK \u2014 the `mannequin` discipline takes the deployment title you simply created:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> openai <span class=\"hljs-keyword\">import<\/span> OpenAI<\/p>\n<p>api_key  = shopper.accounts.list_keys(RESOURCE_GROUP, ACCOUNT_NAME).key1<br \/>\nendpoint = <span class=\"hljs-string\">f&#8221;https:\/\/<span class=\"hljs-subst\">{ACCOUNT_NAME}<\/span>.companies.ai.azure.com\/openai\/v1&#8243;<\/span><\/p>\n<p>openai_client = OpenAI(base_url=endpoint, api_key=api_key)<\/p>\n<p>completion = openai_client.chat.completions.create(<br \/>\n    mannequin=deployment.title,<br \/>\n    messages=[{<span class=\"hljs-string\">&#8220;role&#8221;<\/span>: <span class=\"hljs-string\">&#8220;user&#8221;<\/span>, <span class=\"hljs-string\">&#8220;content&#8221;<\/span>: <span class=\"hljs-string\">&#8220;What is the capital of France?&#8221;<\/span>}],<br \/>\n)<\/p>\n<p><span class=\"hljs-built_in\">print<\/span>(completion.decisions[<span class=\"hljs-number\">0<\/span>].message)<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tUse It in an Agent<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>A chat-completions mannequin from the Assortment slots into Foundry Brokers as an admin-connected mannequin and is callable by way of the Foundry Responses API with the identical OpenAI SDK \u2014 similar auth, similar endpoint, similar observability.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhat&#8217;s Accessible At present<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Accessible now in preview: the Hugging Face Assortment within the Microsoft Foundry Mannequin Catalog \u2014 1000&#8217;s of fashions throughout each modality, refreshed weekly, deployable onto Foundry Managed Compute with NVIDIA A100, NVIDIA H100, or AMD MI300X accelerators in International and Knowledge Zone scopes, behind a unified Foundry endpoint with Playground assist, first-class Azure Monitor metrics, per-deployment billing tags, and curated runtime upgrades and CVE patching utilized routinely to your deployments.<\/p>\n<p>Join the preview: kinds.cloud.microsoft\/r\/8Jnx1LALLA<\/p>\n<p>On the roadmap: broader protection of the Hugging Face ecosystem, further accelerator households, and Convey Your Personal Weights for fine-tuned and proprietary variants deployed by way of the identical templates and governance as Assortment fashions.<\/p>\n<p>Hugging Face is the place open fashions are revealed and found. Microsoft Foundry is the place enterprises operationalize them \u2014 on curated, license-screened, security-screened weights hosted in Azure; on community-built and CVE-scanned runtimes; behind a single endpoint with enterprise identification, networking, observability, and agent integration on prime. **The breadth of the open-source ecosystem, with the operational layer Microsoft runs beneath.**For a deep dive on Foundry Managed Compute \u2014 pricing, accelerator SKUs, knowledge residency, enterprise readiness, observability, and the total Responses API + reminiscence sample \u2014 see the Managed Compute launch weblog.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/microsoft\/foundry-managed-compute\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>At Microsoft Construct 2026, we introduced Foundry Managed Compute and Hugging Face fashions on Foundry \u2014 a curated catalog of open-weight fashions from the Hugging Face ecosystem, refreshed weekly, deployable in a single click on onto Foundry Managed Compute. Weights are pre-staged in Azure, runtimes are constructed and scanned by Microsoft, and each mannequin within [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2002,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[1478,2304,185,2303,806,293],"class_list":["post-2000","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-compute","tag-face","tag-foundry","tag-hugging","tag-managed","tag-models"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Hugging Face Fashions on Foundry Managed Compute - Future News 24<\/title>\n<meta name=\"description\" content=\"A Blog post by Microsoft on Hugging Face\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Hugging Face Fashions on Foundry Managed Compute - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A Blog post by Microsoft on Hugging Face\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-07T15:20:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-07T19:59:04+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Hugging Face Fashions on Foundry Managed Compute\",\"datePublished\":\"2026-07-07T15:20:00+00:00\",\"dateModified\":\"2026-07-07T19:59:04+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/\"},\"wordCount\":2234,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/688797c1629e0ef013c2556c\\\/Dn2wmfWGM83ZhuQ-C5jPo.png\",\"keywords\":[\"compute\",\"Face\",\"Foundry\",\"Hugging\",\"Managed\",\"Models\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/\",\"name\":\"Hugging Face Fashions on Foundry Managed Compute - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/688797c1629e0ef013c2556c\\\/Dn2wmfWGM83ZhuQ-C5jPo.png\",\"datePublished\":\"2026-07-07T15:20:00+00:00\",\"dateModified\":\"2026-07-07T19:59:04+00:00\",\"description\":\"A Blog post by Microsoft on Hugging Face\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#primaryimage\",\"url\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/688797c1629e0ef013c2556c\\\/Dn2wmfWGM83ZhuQ-C5jPo.png\",\"contentUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/688797c1629e0ef013c2556c\\\/Dn2wmfWGM83ZhuQ-C5jPo.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/07\\\/foundry-managed-compute\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Hugging Face Fashions on Foundry Managed Compute\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Hugging Face Fashions on Foundry Managed Compute - Future News 24","description":"A Blog post by Microsoft on Hugging Face","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/","og_locale":"en_US","og_type":"article","og_title":"Hugging Face Fashions on Foundry Managed Compute - Future News 24","og_description":"A Blog post by Microsoft on Hugging Face","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/","og_site_name":"Future News 24","article_published_time":"2026-07-07T15:20:00+00:00","article_modified_time":"2026-07-07T19:59:04+00:00","og_image":[{"url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Hugging Face Fashions on Foundry Managed Compute","datePublished":"2026-07-07T15:20:00+00:00","dateModified":"2026-07-07T19:59:04+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/"},"wordCount":2234,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png","keywords":["compute","Face","Foundry","Hugging","Managed","Models"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/","name":"Hugging Face Fashions on Foundry Managed Compute - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png","datePublished":"2026-07-07T15:20:00+00:00","dateModified":"2026-07-07T19:59:04+00:00","description":"A Blog post by Microsoft on Hugging Face","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#primaryimage","url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png","contentUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/688797c1629e0ef013c2556c\/Dn2wmfWGM83ZhuQ-C5jPo.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/07\/foundry-managed-compute\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Hugging Face Fashions on Foundry Managed Compute"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2000","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2000"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2000\/revisions"}],"predecessor-version":[{"id":2001,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2000\/revisions\/2001"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2002"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2000"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2000"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2000"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}