Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home AI Platforms & Apps

Run Native Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA

Future News 24 by Future News 24
August 11, 2026
in AI Platforms & Apps
0 0
0
Run Native Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


Meta returns to the open supply ecosystem with the discharge of Muse Glimmer, a 30B open-weight dense mannequin with a 120K+ context window constructed for native AI agentic work.  

Optimized to run throughout a variety of NVIDIA edge, desktop, and workstation AI platforms, Muse Glimmer delivers 20K tokens/sec on a single GPU, enabling always-on brokers to course of knowledge regionally and execute advanced, multi-step workflows. 

Constructed for long-running brokers, not simply conversations 

Most LLMs are optimized for chat, prioritizing single-turn interactions and quick time to first token—however agentic workloads demand a special method. An agent scaffolding a software program undertaking, revising documentation, or managing a information base could execute a number of sequential instrument calls in a single session, whereas requiring a degree of reliability, consistency, long-context coherence, and sustained throughput that chat-first fashions aren’t constructed for. 

Muse Glimmer makes use of a dense structure that prompts each parameter for every token it processes, with no routing, knowledgeable choice, or variance throughout token pathways. In consequence, it excels at agentic workloads that demand dependable instruction following, long-context coherence, predictable latency, and fewer failure modes. 

Side-by-side diagram, showing a dense model activating all 30B parameters per token versus an example of an MoE model routing to 2 of 7 experts Side-by-side diagram, showing a dense model activating all 30B parameters per token versus an example of an MoE model routing to 2 of 7 experts 
Determine 1. Overview of the Muse Glimmer dense mannequin structure in comparison with MoE structure 

Privateness by design throughout native {hardware} 

Agentic workflows involving private information, communications, credentials, and proprietary paperwork require inference that by no means leaves the machine. Muse Glimmer hits an optimum steadiness. It’s massive sufficient for advanced multi-step reasoning, however sufficiently small to suit throughout the VRAM of a single NVIDIA GPU, without having for mannequin sharding, CPU offloading, or utilizing exterior endpoints.  

NVIDIA Tensor Core structure accelerates precisely this compute sample, enabling real-time agentic inference absolutely on system at full context size. 

NVIDIA GeForce RTX 5090 pairs 32 GB of VRAM with fifth-generation Tensor Cores, bringing Muse Glimmer to native developer gadgets, conserving proprietary code on system, and eliminating per-token inference price. 

NVIDIA DGX Spark brings workstation-class efficiency and enterprise agentic pipelines right into a compact system. NVIDIA NVLink supplies high-speed entry to reminiscence, and NVIDIA NIM containers make native Muse Glimmer deployment a one-command operation.  

NVIDIA DGX Station brings rack-scale Blackwell Extremely compute to on-prem enterprise environments for groups working beneath air-gap mandates or compliance frameworks the place cloud inference isn’t an choice.  

NVIDIA Jetson extends native Muse Glimmer inference to the sting, enabling robotics, industrial automation, and embedded programs, the place community isolation is a tough requirement, and each inference choice should occur on the level of motion. 

Optimized Muse Glimmer Efficiency on NVIDIA Blackwell Extremely 

On NVIDIA Blackwell Extremely, Muse Glimmer delivers over 20K tokens/sec/GPU at BF16/NVF4 precision, with the throughput-interactivity curve displaying the 30B dense structure sustaining excessive concurrency with out the routing overhead of MoE fashions. 

A single Blackwell Extremely handles the total mannequin in VRAM with headroom for giant KV cache buffers, making it well-suited for top throughput and low latency that builders have to run always-on brokers solely on native infrastructure. 

Muse Glimmer on NVIDIA Blackwell Ultra via vLLM, delivering over 20K tokens/gpu. Muse Glimmer on NVIDIA Blackwell Ultra via vLLM, delivering over 20K tokens/gpu. 
Determine 2. Muse Glimmer efficiency on NVIDIA Blackwell Extremely throughput at BF16 precision 

Constructing and fine-tuning agentic use instances 

Run NVIDIA NemoClaw in a safe OpenShell atmosphere to create long-running private assistants powered for duties like code technology, private assistant, autonomous help, and extra.  

Architecture diagram showing the NemoClaw OpenClaw agent harness connected to Muse Glimmer via vLLM on DGX Spark. Architecture diagram showing the NemoClaw OpenClaw agent harness connected to Muse Glimmer via vLLM on DGX Spark. 
Determine 3. Muse Glimmer operating regionally with the NemoClaw agent harness in a ruled sandbox, served by vLLM on DGX Spark 

Builders can additional post-train the mannequin utilizing the NVIDIA NeMo AutoModel with high-throughput effectivity, which is a fine-tuning library for native Hugging Face checkpoint help with no mannequin conversion necessities.  

It allows full SFT and LoRA fine-tuning out of the field, optimized for fast experimentation on NVIDIA GPUs, together with DGX Spark. Builders may carry out reinforcement studying with NeMo RL, with pattern recipes and reference accuracy validation curves.  

Video 1. Run NeMoClaw and vLLM on DGX Spark

Versatile deployment paths for Muse Glimmer 

NVIDIA helps a number of inference stacks to fulfill quite a lot of developer wants. 

SGLang and vLLM present open-source inference recipes for builders who require deeper management over efficiency on the NVIDIA accelerated platform. 

It’s additionally accessible as a downloadable NVIDIA NIM, a prebuilt, optimized inference container that auto-selects runtime configuration and serving setup, so groups can give attention to constructing and scaling brokers. 

Get began with Muse Glimmer and native AI brokers  

To get began, obtain Muse Glimmer weights from HuggingFace and deploy utilizing the inference recipes above, or pull the downloadable NIM for a production-ready container on any NVIDIA GPU-accelerated platform. To name a hosted endpoint immediately, strive it on construct.nvidia.com. For edge deployments on Jetson, discover the Jetson AI Lab.



Source link

Tags: AgenticGlimmerLocalMetasMuseNVIDIArunWorkflows
Previous Post

International New Unicorn Counts In The First Half Of 2026 Have Already Surpassed 2025’s Totals

Next Post

Most Enterprise AI Isn’t Safe. Right here’s What Companies Can Do – Unite.AI

Next Post
Most Enterprise AI Isn’t Safe. Right here’s What Companies Can Do – Unite.AI

Most Enterprise AI Isn’t Safe. Right here’s What Companies Can Do – Unite.AI

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb