{"id":1443,"date":"2026-06-24T16:00:00","date_gmt":"2026-06-24T16:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/"},"modified":"2026-06-24T23:59:31","modified_gmt":"2026-06-24T23:59:31","slug":"accelerating-fine-tuning-nvidia-nemo-automodel","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/","title":{"rendered":"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\nHuggingFace Transformers has turn into the inspiration of the open-source AI ecosystem, and the current Transformers v5 launch strengthened it with first-class help for Combination-of-Consultants (MoE) fashions, now the dominant structure for frontier fashions. v5 ships the MoE foundations: professional backends, dynamic weight loading, and distributed execution that make MoE extensible and simple to construct on. <\/p>\n<p>NVIDIA NeMo AutoModel is an open library a part of the NVIDIA NeMo framework for constructing customized generative AI fashions at scale. NeMo AutoModel builds cleanly on prime of v5, including Professional Parallelism, DeepEP fused all-to-all dispatch, and TransformerEngine kernels, and it leans on v5&#8217;s dynamic weight loading to deliver these optimizations to a broad and rising set of mannequin households. The payoff is 3.4-3.7x larger coaching throughput and 29-32% much less GPU reminiscence on fine-tuning MoE fashions than native Transformers v5, utilizing the identical from_pretrained() API: a single import line, with no different code adjustments. <\/p>\n<p>This weblog particulars how this mix works and the way customers can fine-tune MoE fashions sooner with out altering their APIs.<\/p>\n<p>The rise of MoE fashions has launched new challenges to environment friendly coaching: Routing tokens throughout a whole lot of consultants, fusing professional matmuls right into a single kernel, sharding weights throughout GPUs, and overlapping communication with computation all require infrastructure past what a general-purpose library gives out of the field.<\/p>\n<p>Transformers v5 (\u201cv5\u201d) launched first-class MoE help corresponding to professional backends, dynamic weight loading, and tensor parallel plans for distributed execution. As well as, v5 made distributed coaching first-class by integrating PyTorch&#8217;s DeviceMesh instantly into from_pretrained().<\/p>\n<p>NeMo AutoModel builds on prime of v5 by subclassing AutoModelForCausalLM, and including Professional Parallelism (EP), DeepEP fused all-to-all dispatch, and TransformerEngine kernels. DeepEP is the piece v5 does not have but: it overlaps communication with professional compute. And since NeMo AutoModel rides v5&#8217;s reversible weight conversion to load every mannequin, it will probably focus its engineering on these reusable core ops as an alternative of per-model checkpoint plumbing, whereas save_pretrained() nonetheless emits commonplace HF checkpoints that instruments like vLLM and SGLang can load. <\/p>\n<p>The following part walks by means of how the 2 work collectively and the efficiency good points we measured, from full fine-tuning NVIDIA Nemotron 3 Extremely 550B A55B throughout 16 nodes all the way down to single-node fashions corresponding to Qwen3-30B-A3B and Nemotron 3 Nano 30B A3B.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tNeMo AutoModel: Similar API, Extra Efficiency<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Considered one of NeMo AutoModel&#8217;s targets is API compatibility with HuggingFace Transformers to allow open-source neighborhood. NeMoAutoModelForCausalLM subclasses AutoModelForCausalLM, so any code that works with HF fashions works with AutoModel too.<\/p>\n<p>Here is what loading a mannequin seems to be like in each. Solely the import adjustments:<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/VTPq2Wp-RrEcP1eGUJxao.png\" alt=\"nemo_and_hf\"\/><\/p>\n<p>That single import does lots of work. For well-liked MoE architectures like Qwen3, NVIDIA Nemotron, GPT-OSS, and DeepSeek V3, NeMo AutoModel ships hand-tuned implementations with TransformerEngine consideration, fused linear layers, and customized professional kernels. For the whole lot else, it falls again to vanilla HF whereas nonetheless making use of optimizations like Liger kernel patching, amongst others. And whichever path it takes, the ensuing mannequin is able to scale: move a device_mesh and you&#8217;ve got multi-GPU coaching with out additional rewrites.<\/p>\n<p>The place NeMo AutoModel actually shines is scaling MoE fashions to multi-GPU coaching. To coach Nemotron 3 Nano 30B A3B with Professional Parallelism throughout 8 GPUs, one provides the distributed mesh configuration:<\/p>\n<p><span class=\"hljs-keyword\">import<\/span> os<br \/>\n<span class=\"hljs-keyword\">import<\/span> torch<br \/>\n<span class=\"hljs-keyword\">import<\/span> torch.distributed <span class=\"hljs-keyword\">as<\/span> dist<br \/>\n<span class=\"hljs-keyword\">from<\/span> nemo_automodel <span class=\"hljs-keyword\">import<\/span> NeMoAutoModelForCausalLM<br \/>\n<span class=\"hljs-keyword\">from<\/span> nemo_automodel.recipes._dist_utils <span class=\"hljs-keyword\">import<\/span> create_distributed_setup_from_config<\/p>\n<p>dist.init_process_group(backend=<span class=\"hljs-string\">&#8220;nccl&#8221;<\/span>)<br \/>\ntorch.manual_seed(<span class=\"hljs-number\">0<\/span>)<br \/>\ntorch.cuda.set_device(<span class=\"hljs-built_in\">int<\/span>(os.environ.get(<span class=\"hljs-string\">&#8220;LOCAL_RANK&#8221;<\/span>, <span class=\"hljs-number\">0<\/span>)))<\/p>\n<p>dist_setup = create_distributed_setup_from_config(<br \/>\n    {<br \/>\n        <span class=\"hljs-string\">&#8220;technique&#8221;<\/span>: <span class=\"hljs-string\">&#8220;fsdp2&#8221;<\/span>,<br \/>\n        <span class=\"hljs-string\">&#8220;ep_size&#8221;<\/span>: <span class=\"hljs-number\">8<\/span>,<br \/>\n    },<br \/>\n)<\/p>\n<p>mannequin = NeMoAutoModelForCausalLM.from_pretrained(<br \/>\n    <span class=\"hljs-string\">&#8220;nvidia\/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16&#8221;<\/span>,<br \/>\n    dtype=torch.bfloat16,<br \/>\n    distributed_setup=dist_setup,<br \/>\n)<\/p>\n<p>dist.destroy_process_group()<\/p>\n<p>This offers velocity, scalability and memory-optimizations with FSDP2,  Professional Parallelism,  TransformerEngine kernels and DeepEP dispatch, all from a from_pretrained() name.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tEfficiency Comparability<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>We evaluated NeMo AutoModel in two regimes: full fine-tuning a frontier-scale 550B mannequin throughout 16 nodes, and coaching two 30B MoE fashions on a single node. The 550B outcome reveals why Professional Parallelism is important at scale; the 30B outcomes quantify the per-GPU speedup over Transformers v5.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tNemotron 3 Extremely 550B A55B (full fine-tune, multi-node)<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Nemotron 3 Extremely 550B A55B is a 550B-parameter hybrid mannequin transport with Mamba2, LatentMoE, and Multi-Token Prediction (MTP). We benchmark a full fine-tune: each parameter is up to date and the Adam optimizer state is materialized, which at this scale spans 16 H100 nodes (128 GPUs).<\/p>\n<p>Methodology:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Parameter<br \/>\nWorth<\/p>\n<p>{Hardware}<br \/>\n16x H100 80GB (128 GPUs)<\/p>\n<p>Professional Parallelism<br \/>\nEP=64<\/p>\n<p>Native batch dimension<br \/>\n2<\/p>\n<p>Sequence size<br \/>\n4,096<\/p>\n<p>Options<br \/>\nMTP, activation checkpointing, fused linear cross-entropy<\/p>\n<p>Kernels<br \/>\nDeepEP dispatch + torch_mm consultants + TransformerEngine<\/p>\n<\/div>\n<div class=\"max-w-full overflow-auto\">\n<p>Metric<br \/>\nNeMo AutoModel (EP=64)<\/p>\n<p>TPS\/GPU (avg)<br \/>\n815<\/p>\n<p>TFLOP\/s\/GPU<br \/>\n~293<\/p>\n<p>Peak Reminiscence<br \/>\n58.2 GiB<\/p>\n<\/div>\n<p>Why there is no such thing as a Transformers v5 column. Transformers v5 runs out of reminiscence at this scale, so there is no such thing as a v5 quantity to report right here. AutoModel&#8217;s Professional Parallelism shards the consultants throughout GPUs to deliver the footprint inside finances, which is what lets the complete fine-tune run. The 30B comparisons under present the identical benefit the place v5 matches.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tSingle-node 30B MoE benchmarks<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>We benchmarked three approaches on a single node with 8x H100 80GB GPUs: HF Transformers v4 (hub code), HF Transformers v5 (with greatest obtainable optimizations), and NeMo AutoModel (EP=8 + customized kernels).<\/p>\n<p>Methodology:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Parameter<br \/>\nWorth<\/p>\n<p>{Hardware}<br \/>\n8x H100 80GB (single node)<\/p>\n<p>Sequence size<br \/>\n4,096<\/p>\n<p>Native batch dimension<br \/>\n1<\/p>\n<\/div>\n<p>A notice on the routing gate. The NeMo AutoModel numbers under use a balanced routing gate, which forces tokens to be distributed uniformly throughout consultants. This emulates the perfect working level an MoE is skilled towards: a well-trained mannequin&#8217;s load-balancing loss drives professional utilization to near-uniform, so balanced routing displays the steady-state an actual workload converges to (and removes the straggler noise that random dummy tokens in any other case inject into professional parallelism). v4\/v5 run their native router on the identical dummy tokens. The balanced gate subsequently measures NeMo AutoModel at its goal MoE working level, and the v4\/v5 columns mirror their out-of-the-box conduct.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/rbCVgV6a18c4UcDsiWfZN.png\" alt=\"nemo_automodel_blog_chart_mockup_v5\"\/><\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tQwen3-30B-A3B<br \/>\n\t<\/span><br \/>\n<\/h3>\n<div class=\"max-w-full overflow-auto\">\n<p>Metric<br \/>\nv4<br \/>\nv5 (FA2 + grouped_mm)<br \/>\nNeMo AutoModel (EP=8)<br \/>\nv5 \u2192 NeMo AutoModel<\/p>\n<p>TPS\/GPU (avg)<br \/>\nimpasse<br \/>\n3,075<br \/>\n11,340<br \/>\n3.69x<\/p>\n<p>Peak Reminiscence<br \/>\n\u2014<br \/>\n68.2 GiB<br \/>\n48.1 GiB<br \/>\n-29%<\/p>\n<p>Avg Ahead+Loss<br \/>\n\u2014<br \/>\n582 ms<br \/>\n194 ms<br \/>\n3.00x<\/p>\n<p>Avg Backward<br \/>\n\u2014<br \/>\n758 ms<br \/>\n178 ms<br \/>\n4.26x<\/p>\n<\/div>\n<p>Why v4 deadlocks: Transformers v4 shops Qwen3 MoE consultants as a ModuleList of 128 particular person MLP modules, every individually FSDP-wrapped. The ahead move makes use of a data-dependent loop that solely iterates consultants that acquired tokens. With completely different knowledge per rank, completely different ranks skip completely different consultants, inflicting mismatched FSDP AllGather\/ReduceScatter collectives and an indefinite hold. Transformers v5 fixes this by storing consultants as fused 3D parameter tensors (no per-expert modules, no per-expert FSDP collectives).<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tNemotron 3 Nano 30B A3B<br \/>\n\t<\/span><br \/>\n<\/h3>\n<div class=\"max-w-full overflow-auto\">\n<p>Metric<br \/>\nv4 (hub code)<br \/>\nv5 (FA2 + grouped_mm + Mamba CUDA)<br \/>\nNeMo AutoModel (EP=8)<br \/>\nv5 \u2192 NeMo AutoModel<\/p>\n<p>TPS\/GPU (avg)<br \/>\n1,807<br \/>\n4,583<br \/>\n15,421<br \/>\n3.36x<\/p>\n<p>Peak Reminiscence<br \/>\n61.9 GiB<br \/>\n62.1 GiB<br \/>\n42.5 GiB<br \/>\n-32%<\/p>\n<p>Avg Ahead+Loss<br \/>\n1,024 ms<br \/>\n283 ms<br \/>\n109 ms<br \/>\n2.60x<\/p>\n<p>Avg Backward<br \/>\n1,246 ms<br \/>\n611 ms<br \/>\n157 ms<br \/>\n3.89x<\/p>\n<\/div>\n<p>v4 config: trust_remote_code=True (NVIDIA&#8217;s hub modeling code). The hub code&#8217;s professional loop is FSDP-safe (iterates all consultants no matter token project), so it does not impasse like Qwen3 v4.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe place the speedup comes from<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>The three.4-3.7x speedup from NeMo AutoModel over Transformers v5 comes from three sources:<\/p>\n<p>Professional Parallelism reduces reminiscence strain. EP=8 distributes professional weights throughout GPUs, reducing the per-GPU MoE footprint by 8x. For Qwen3, this drops peak reminiscence from 68.2 GiB to 48.1 GiB (-29%). For Nemotron Nano, it drops from 62.1 GiB to 42.5 GiB (-32%), liberating headroom for bigger batch sizes or longer sequences.  <\/p>\n<p>DeepEP fuses communication with computation. As a substitute of separate AllGather\/ReduceScatter collectives for professional routing, DeepEP fuses token dispatch and combines into optimized GPU kernels, overlapping communication with professional computation.  <\/p>\n<p>TransformerEngine kernels speed up core operations. TE&#8217;s fused consideration, linear layers, and RMSNorm implementations present constant speedups over their PyTorch\/Flash Consideration equivalents throughout all layer varieties, not simply MoE layers.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tTransformers v5 Options Leveraged by HuggingFace AutoModel<br \/>\n\t<\/span><br \/>\n<\/h2>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tProfessional Backends<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Some of the impactful options in Transformers v5 is the experts_implementation parameter, which incorporates three professional backends:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Backend<br \/>\nDescription<br \/>\nGreatest for<\/p>\n<p>keen<br \/>\nFor-loop over chosen consultants<br \/>\nDebugging, compatibility, and correctness. Additionally obtainable for v4.<\/p>\n<p>batched_mm<br \/>\nDuplicates professional params, single batched GEMM through torch.bmm<br \/>\nSmall inputs, quick with torch.compile. Added for v5<\/p>\n<p>grouped_mm<br \/>\nOrders tokens by professional, single grouped GEMM through torch.nn.useful.grouped_mm<br \/>\nCoaching (reminiscence environment friendly, no param duplication). Added for v5.<\/p>\n<\/div>\n<p>The grouped_mm backend is the important thing coaching optimization: as an alternative of looping over consultants one after the other, it kinds tokens by their assigned professional and executes a single fused grouped matrix multiplication.<\/p>\n<p>NeMo AutoModel takes this additional. For fashions with customized implementations, it makes use of DeepEP fused all-to-all dispatch mixed with grouped GEMM kernels and TransformerEngine linear layers. The development seems to be like:<\/p>\n<p>v4 (keen for-loop) \u2192 v5 (grouped_mm) \u2192 NeMo AutoModel (DeepEP + GMM + TE)<\/p>\n<p>In NeMo AutoModel, the professional backend is configured by means of BackendConfig:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> nemo_automodel.elements.fashions.frequent.utils <span class=\"hljs-keyword\">import<\/span> BackendConfig<\/p>\n<p>backend = BackendConfig(<br \/>\n    attn=<span class=\"hljs-string\">&#8220;te&#8221;<\/span>,<br \/>\n    linear=<span class=\"hljs-string\">&#8220;te&#8221;<\/span>,<br \/>\n    consultants=<span class=\"hljs-string\">&#8220;torch_mm&#8221;<\/span>,<br \/>\n    dispatcher=<span class=\"hljs-string\">&#8220;deepep&#8221;<\/span>,<br \/>\n)<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tProfessional Parallelism and DeepEP<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Transformers v5 additionally ships an Professional Parallelism path. It shards professional weights throughout GPUs. The GroupedGemmParallel model masses solely every system&#8217;s native consultants, and RouterParallel routes tokens and combines outcomes with an all_reduce. It is neatly constructed on v5&#8217;s present tensor-parallel equipment. Enabling it makes the mannequin&#8217;s tp_plan return its professional plan, so professional parallelism shares the system finances with knowledge parallelism (ep \u00d7 dp = world_size). For the single-node 30B benchmarks right here, we discovered plain data-parallel v5 (dp=8, ep=1) to be the quickest v5 configuration, so that is the v5 setup we report.<\/p>\n<p>NeMo AutoModel takes a complementary method tuned for multi-GPU MoE coaching. It makes EP its personal parallelism dimension, a devoted moe_mesh alongside (slightly than carved from) the data-parallel mesh, utilizing PyTorch&#8217;s DTensor with Shard(0). As a result of the professional mesh is orthogonal to knowledge parallelism, the 2 compose on the identical gadgets. On 8 GPUs NeMo AutoModel runs ep=8 and dp=8 collectively, so each GPU trains by itself knowledge shard whereas holding just one\/8 of the consultants. Professional weights are bodily sharded throughout GPUs alongside the professional dimension.<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> torch.distributed.tensor <span class=\"hljs-keyword\">import<\/span> Shard, distribute_tensor<\/p>\n<p>distribute_tensor(param, device_mesh, [Shard(<span class=\"hljs-number\">0<\/span>)])<\/p>\n<p>With ep_size=8 on 8 GPUs, every GPU holds just one\/8 of the professional parameters. For a mannequin like Nemotron-3-Nano-30B-A3B with ~55 GiB of professional weights, EP reduces the per-GPU professional footprint from ~55 GiB to ~6.8 GiB, making coaching potential the place FSDP-only approaches run out of reminiscence.<\/p>\n<p>On prime of EP, NeMo AutoModel integrates DeepEP that fuses the token routing into optimized GPU kernels, and delivers vital speedups when mixed with grouped GEMM for grouped professional computation. In our large-scale MoE benchmarks, DeepEP + grouped GEMM diminished price per iteration by 47% on the complete DeepSeek V3 671B mannequin in comparison with all-gather + looped professional baselines.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tDynamic Weight Loading<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Transformers v5 additionally launched a dynamic weight loading system by means of WeightConverter and WeightRenaming. This permits MoE checkpoint to be saved in fused 3D tensors for extra environment friendly execution. The WeightConverter applies composable operations to remodel checkpoint tensors on-the-fly throughout from_pretrained().<\/p>\n<p>NeMo AutoModel is a direct client of this v5 API. Over 20 mannequin varieties use this mechanism by means of MODELS_REQUIRING_TENSOR_MERGING, together with Mixtral, Qwen2 MoE, Qwen3 MoE, DeepSeek V2\/V3, OLMoE, and extra. The conversions are totally reversible: save_pretrained() produces commonplace HF-format checkpoints that any downstream instrument can load.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tGetting Began<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>To strive NeMo AutoModel, please go to our official documentation web page to get began.<\/p>\n<p>For extra particulars, see:<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tConclusion<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>NVIDIA NeMo AutoModel is the pure subsequent step for HuggingFace customers scaling up mannequin coaching. By constructing instantly on Transformers v5, AutoModel gives a zero-friction improve path: change one import line and get a mannequin occasion that&#8217;s greater than 3 times as quick.<\/p>\n<p>On Qwen3-30B-A3B and Nemotron 3 Nano 30B-A3B, this delivers 3.4-3.7x larger coaching throughput with 29-32% much less GPU reminiscence in comparison with the very best Transformers v5 configuration. And since true Professional Parallelism shards consultants throughout GPUs, the identical path scales as much as full fine-tuning a 550B mannequin like Nemotron 3 Extremely throughout 16 nodes, the regime the place Professional Parallelism turns into important to suit the mannequin in reminiscence. As a result of NeMo AutoModel checkpoints are commonplace HF-format safetensors, you&#8217;ll be able to deploy them on inference frameworks like vLLM and SGLang.<\/p>\n<p>The code, configs, and benchmark scripts are all obtainable within the NeMo AutoModel repository.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tAcknowledgements<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Core contributors to this work, listed alphabetically by final title: Adil Asif, Hemil Desai, Alexandros Koumparoulis, and Huiying Li.  <\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/nvidia\/accelerating-fine-tuning-nvidia-nemo-automodel\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>HuggingFace Transformers has turn into the inspiration of the open-source AI ecosystem, and the current Transformers v5 launch strengthened it with first-class help for Combination-of-Consultants (MoE) fashions, now the dominant structure for frontier fashions. v5 ships the MoE foundations: professional backends, dynamic weight loading, and distributed execution that make MoE extensible and simple to construct [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1445,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[1913,1916,1421,1915,81,1914],"class_list":["post-1443","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-accelerating","tag-automodel","tag-finetuning","tag-nemo","tag-nvidia","tag-transformers"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel - Future News 24<\/title>\n<meta name=\"description\" content=\"A Blog post by NVIDIA on Hugging Face\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A Blog post by NVIDIA on Hugging Face\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-24T16:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-24T23:59:31+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel\",\"datePublished\":\"2026-06-24T16:00:00+00:00\",\"dateModified\":\"2026-06-24T23:59:31+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/\"},\"wordCount\":2177,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/690d0a6c2c5acfe0e1f4777d\\\/1N4GjIYBsZ6RCReRx_qBB.png\",\"keywords\":[\"Accelerating\",\"AutoModel\",\"FineTuning\",\"NeMo\",\"NVIDIA\",\"Transformers\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/\",\"name\":\"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/690d0a6c2c5acfe0e1f4777d\\\/1N4GjIYBsZ6RCReRx_qBB.png\",\"datePublished\":\"2026-06-24T16:00:00+00:00\",\"dateModified\":\"2026-06-24T23:59:31+00:00\",\"description\":\"A Blog post by NVIDIA on Hugging Face\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#primaryimage\",\"url\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/690d0a6c2c5acfe0e1f4777d\\\/1N4GjIYBsZ6RCReRx_qBB.png\",\"contentUrl\":\"https:\\\/\\\/cdn-uploads.huggingface.co\\\/production\\\/uploads\\\/690d0a6c2c5acfe0e1f4777d\\\/1N4GjIYBsZ6RCReRx_qBB.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-fine-tuning-nvidia-nemo-automodel\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel - Future News 24","description":"A Blog post by NVIDIA on Hugging Face","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/","og_locale":"en_US","og_type":"article","og_title":"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel - Future News 24","og_description":"A Blog post by NVIDIA on Hugging Face","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/","og_site_name":"Future News 24","article_published_time":"2026-06-24T16:00:00+00:00","article_modified_time":"2026-06-24T23:59:31+00:00","og_image":[{"url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel","datePublished":"2026-06-24T16:00:00+00:00","dateModified":"2026-06-24T23:59:31+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/"},"wordCount":2177,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png","keywords":["Accelerating","AutoModel","FineTuning","NeMo","NVIDIA","Transformers"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/","name":"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png","datePublished":"2026-06-24T16:00:00+00:00","dateModified":"2026-06-24T23:59:31+00:00","description":"A Blog post by NVIDIA on Hugging Face","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#primaryimage","url":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png","contentUrl":"https:\/\/cdn-uploads.huggingface.co\/production\/uploads\/690d0a6c2c5acfe0e1f4777d\/1N4GjIYBsZ6RCReRx_qBB.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-fine-tuning-nvidia-nemo-automodel\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Accelerating Transformers Nice-Tuning with NVIDIA NeMo AutoModel"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1443","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1443"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1443\/revisions"}],"predecessor-version":[{"id":1444,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1443\/revisions\/1444"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1445"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1443"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1443"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1443"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}