Meta returns to the open supply ecosystem with the discharge of Muse Glimmer, a 30B open-weight dense mannequin with a 120K+ context window constructed for native AI agentic work.
Optimized to run throughout a variety of NVIDIA edge, desktop, and workstation AI platforms, Muse Glimmer delivers 20K tokens/sec on a single GPU, enabling always-on brokers to course of knowledge regionally and execute advanced, multi-step workflows.
Constructed for long-running brokers, not simply conversations
Most LLMs are optimized for chat, prioritizing single-turn interactions and quick time to first token—however agentic workloads demand a special method. An agent scaffolding a software program undertaking, revising documentation, or managing a information base could execute a number of sequential instrument calls in a single session, whereas requiring a degree of reliability, consistency, long-context coherence, and sustained throughput that chat-first fashions aren’t constructed for.
Muse Glimmer makes use of a dense structure that prompts each parameter for every token it processes, with no routing, knowledgeable choice, or variance throughout token pathways. In consequence, it excels at agentic workloads that demand dependable instruction following, long-context coherence, predictable latency, and fewer failure modes.


Privateness by design throughout native {hardware}
Agentic workflows involving private information, communications, credentials, and proprietary paperwork require inference that by no means leaves the machine. Muse Glimmer hits an optimum steadiness. It’s massive sufficient for advanced multi-step reasoning, however sufficiently small to suit throughout the VRAM of a single NVIDIA GPU, without having for mannequin sharding, CPU offloading, or utilizing exterior endpoints.
NVIDIA Tensor Core structure accelerates precisely this compute sample, enabling real-time agentic inference absolutely on system at full context size.
NVIDIA GeForce RTX 5090 pairs 32 GB of VRAM with fifth-generation Tensor Cores, bringing Muse Glimmer to native developer gadgets, conserving proprietary code on system, and eliminating per-token inference price.
NVIDIA DGX Spark brings workstation-class efficiency and enterprise agentic pipelines right into a compact system. NVIDIA NVLink supplies high-speed entry to reminiscence, and NVIDIA NIM containers make native Muse Glimmer deployment a one-command operation.
NVIDIA DGX Station brings rack-scale Blackwell Extremely compute to on-prem enterprise environments for groups working beneath air-gap mandates or compliance frameworks the place cloud inference isn’t an choice.
NVIDIA Jetson extends native Muse Glimmer inference to the sting, enabling robotics, industrial automation, and embedded programs, the place community isolation is a tough requirement, and each inference choice should occur on the level of motion.
Optimized Muse Glimmer Efficiency on NVIDIA Blackwell Extremely
On NVIDIA Blackwell Extremely, Muse Glimmer delivers over 20K tokens/sec/GPU at BF16/NVF4 precision, with the throughput-interactivity curve displaying the 30B dense structure sustaining excessive concurrency with out the routing overhead of MoE fashions.
A single Blackwell Extremely handles the total mannequin in VRAM with headroom for giant KV cache buffers, making it well-suited for top throughput and low latency that builders have to run always-on brokers solely on native infrastructure.


Constructing and fine-tuning agentic use instances
Run NVIDIA NemoClaw in a safe OpenShell atmosphere to create long-running private assistants powered for duties like code technology, private assistant, autonomous help, and extra.


Builders can additional post-train the mannequin utilizing the NVIDIA NeMo AutoModel with high-throughput effectivity, which is a fine-tuning library for native Hugging Face checkpoint help with no mannequin conversion necessities.
It allows full SFT and LoRA fine-tuning out of the field, optimized for fast experimentation on NVIDIA GPUs, together with DGX Spark. Builders may carry out reinforcement studying with NeMo RL, with pattern recipes and reference accuracy validation curves.
Versatile deployment paths for Muse Glimmer
NVIDIA helps a number of inference stacks to fulfill quite a lot of developer wants.
SGLang and vLLM present open-source inference recipes for builders who require deeper management over efficiency on the NVIDIA accelerated platform.
It’s additionally accessible as a downloadable NVIDIA NIM, a prebuilt, optimized inference container that auto-selects runtime configuration and serving setup, so groups can give attention to constructing and scaling brokers.
Get began with Muse Glimmer and native AI brokers
To get began, obtain Muse Glimmer weights from HuggingFace and deploy utilizing the inference recipes above, or pull the downloadable NIM for a production-ready container on any NVIDIA GPU-accelerated platform. To name a hosted endpoint immediately, strive it on construct.nvidia.com. For edge deployments on Jetson, discover the Jetson AI Lab.

