{"id":68,"date":"2026-06-04T12:59:00","date_gmt":"2026-06-04T12:59:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/"},"modified":"2026-06-04T18:41:18","modified_gmt":"2026-06-04T18:41:18","slug":"fine-tuning-nemotron-35-asr","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/","title":{"rendered":"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\nIntroducing NVIDIA Nemotron 3.5 ASR, streaming multilingual: a 600M-parameter speech-to-text mannequin that transcribes 40 language-locales from a single checkpoint, in actual time, with punctuation and capitalization inbuilt. It&#8217;s the successor of the favored Nemotron 3 ASR mannequin (English solely) which was launched on Hugging Face and as a NIM earlier this 12 months. Since its launch, Nemotron 3 ASR has been validated by impartial benchmarks at Synthetic Evaluation, the place it ranks 2nd in latency amongst all streaming ASR fashions\u2014 with simply 0.07 seconds to last transcript after finish of speech \u2014 and sits within the &#8220;most tasty quadrant&#8221; of the AA-WER Streaming Index vs. Time to Remaining Transcription leaderboard, putting it among the many finest fashions on the mixed accuracy-latency tradeoff. The mannequin makes use of a Cache-Conscious FastConformer-RNNT structure that streams audio with out the redundant recomputation that makes most streaming ASR sluggish \u2014 so that you get low latency and excessive accuracy, not one on the expense of the opposite. Nemotron 3.5 ASR ships as open weights on Hugging Face \u2014 you possibly can examine, fine-tune, and deploy it with out API dependencies or per-call billing. No knowledge leaves your infrastructure except you select. And since it is a sturdy base mannequin, you possibly can fine-tune it on your personal language, area, or accent. The second half of this submit walks by way of precisely how.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe issue with multilingual speech recognition right this moment<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>For those who&#8217;ve ever constructed a product that should transcribe speech, you have in all probability hit considered one of these partitions:<\/p>\n<p>The polyglot tax. You need to assist a number of languages, so that you sew collectively 40 totally different fashions \u2014 or 40 totally different vendor APIs \u2014 every with its personal quirks, latency profile, and billing. Your infrastructure turns into a museum of one-off integrations.<br \/>\nThe streaming-vs-accuracy tradeoff. Actual-time captioning wants low latency, however most &#8220;streaming&#8221; ASR programs faux it by re-processing overlapping home windows of audio again and again. That burns compute and provides delay. Flip down the latency and accuracy falls off a cliff.<br \/>\nThe post-processing pipeline. Uncooked ASR output is usually an unpunctuated, lowercase wall of textual content. You bolt on a second mannequin for punctuation and capitalization, including one more shifting half.<br \/>\nThe &#8220;recognized language&#8221; assumption. Many programs require you to inform them the language up entrance. However what a few customer-support line the place callers change between English and Spanish mid-sentence?<\/p>\n<p>Nemotron 3.5 ASR was constructed to break down all 4 of these issues into one mannequin.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhat it does<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>One mannequin, 40  language-locales. A single 600M-parameter checkpoint transcribes English (US\/GB), Spanish (US\/ES), German, French (FR\/CA), Italian, Arabic, Japanese, Korean, Portuguese (BR\/PT), Russian, Hindi, Turkish, Vietnamese, Dutch, Ukrainian, Polish, Finnish, Mandarin, Czech, Bulgarian, Slovak, Swedish, Croatian, Romanian, Estonian, Danish, Hungarian, Norwegian Bokm\u00e5l, Norwegian Nynorsk, Hebrew, Greek, Lithuanian, Latvian, Maltese, Slovenian, and Thai. No per-language deployment, no model-swapping.<\/p>\n<p>Actual-time streaming, completed proper. The mannequin is constructed on a Cache-Conscious FastConformer encoder. Conventional &#8220;buffered&#8221; streaming re-processes overlapping chunks of audio at each step, doing the identical work many instances over. This mannequin as an alternative caches the encoder&#8217;s inner state and reuses it \u2014 each audio body is processed precisely as soon as, with no overlap. The result&#8217;s dramatically decrease compute and end-to-end latency, with no accuracy penalty.<\/p>\n<p>Punctuation and capitalization, natively. The output is production-ready textual content \u2014 correct casing, commas, durations, query marks \u2014 straight from the mannequin. No separate punctuation-restoration step.<\/p>\n<p>Language conditioning, your alternative. You&#8217;ll be able to run it two methods:<\/p>\n<p>Inform the mannequin the enter language (target_lang=en-US) when you understand it \u2014 sometimes the very best accuracy.<br \/>\nLet the mannequin detect the language (target_lang=auto) when you do not \u2014 the mannequin detects the language and transcribes accordingly.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe way it works (the 2-minute model)<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The mannequin has two fundamental items:<\/p>\n<p>A Cache-Conscious FastConformer encoder (24 layers). FastConformer is an environment friendly evolution of the Conformer structure with linearly scalable consideration. The &#8220;cache-aware&#8221; half is the streaming magic: the encoder retains a cache of its self-attention and convolution activations from earlier frames, in order new audio arrives it solely computes what&#8217;s genuinely new. Nothing is recomputed.  <\/p>\n<p>An RNNT (Recurrent Neural Community Transducer) decoder. RNNT is the workhorse decoder for streaming ASR \u2014 it emits textual content as audio streams in, body by body, which is precisely what you need for reside transcription.<\/p>\n<p>On prime of this, the mannequin provides prompt-based language-ID conditioning: a language sign is fed alongside the audio, which lets one set of weights specialize its output to the goal language \u2014 or, in auto mode, infer the language itself.<\/p>\n<p>It was educated on an enormous speech knowledge spanning all supported languages, utilizing a mix of public and proprietary knowledge normalized to punctuated, properly-cased textual content.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tA knob value figuring out: att_context_size<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Streaming ASR is essentially a tradeoff between how quickly you emit textual content and the way a lot future audio the mannequin will get to &#8220;peek at&#8221; earlier than committing. Nemotron ASR exposes this immediately by way of the eye context measurement:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Consideration Context<br \/>\nChunk Dimension (Latency)<br \/>\nUse Case<\/p>\n<p>[56, 0]<br \/>\n80ms (Extremely-Low)<br \/>\nExtremely low latency Voice Brokers<\/p>\n<p>[56, 1]<br \/>\n160ms (Low)<br \/>\nInteractive Voice Brokers, Conversational AI<\/p>\n<p>[56, 3]<br \/>\n320ms (Balanced)<br \/>\nConversational AI, Dwell caption<\/p>\n<p>[56, 6]<br \/>\n560ms (Medium)<br \/>\nExcessive accuracy with affordable latency<\/p>\n<p>[56, 13]<br \/>\n1.12s (Excessive)<br \/>\nHighest accuracy with excessive latency<\/p>\n<\/div>\n<p>The identical checkpoint covers the entire spectrum \u2014 you select the working level at inference time, no retraining required.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tAttempt it in minutes<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The mannequin ships as a NeMo checkpoint. Clone the NeMo department and level the streaming inference script at your audio:<\/p>\n<p>git clone https:\/\/github.com\/NVIDIA-NeMo\/NeMo.git<\/p>\n<p>Transcribe with a recognized language:<\/p>\n<p>python ${NEMO_ROOT}\/examples\/asr\/asr_cache_aware_streaming\/speech_to_text_cache_aware_streaming_infer.py<br \/>\n    model_path=${MODEL_PATH}<br \/>\n    dataset_manifest=${MANIFEST_PATH}<br \/>\n    output_path=${OUTPUT_FOLDER}<br \/>\n    target_lang=es-ES<br \/>\n    att_context_size=&#8221;[56,3]&#8221;<br \/>\n    strip_lang_tags=true<\/p>\n<p>Or let the mannequin detect the language:<\/p>\n<p>python ${NEMO_ROOT}\/examples\/asr\/asr_cache_aware_streaming\/speech_to_text_cache_aware_streaming_infer.py<br \/>\n    model_path=${MODEL_PATH}<br \/>\n    dataset_manifest=${MANIFEST_PATH}<br \/>\n    output_path=${OUTPUT_FOLDER}<br \/>\n    target_lang=auto<br \/>\n    att_context_size=&#8221;[56,3]&#8221;<br \/>\n    strip_lang_tags=true<\/p>\n<p>Audio ought to be mono-channel .wav. The manifest is an ordinary NeMo JSON-lines file:<\/p>\n<p><span class=\"hljs-punctuation\">{<\/span><span class=\"hljs-attr\">&#8220;audio_filepath&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-string\">&#8220;\/path\/to\/clip.wav&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span> <span class=\"hljs-attr\">&#8220;length&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-number\">4.27<\/span><span class=\"hljs-punctuation\">,<\/span> <span class=\"hljs-attr\">&#8220;textual content&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-string\">&#8220;reference transcript&#8221;<\/span><span class=\"hljs-punctuation\">}<\/span><\/p>\n<p>Mannequin robotically predicts language_tag on the finish of every accomplished sentence, i.e. \u201cIt is a check pattern. \u201d. \u201cstrip_lang_tags=True\u201d removes the language tag  for higher readability. <\/p>\n<p>Nemotron 3.5 ASR is powerful out of the field \u2014 nevertheless it was educated on a combination the place some languages have way more knowledge than others. The long-tail locales have headroom, and some hours of in-domain audio plus the suitable recipe closes a shocking quantity of it.<\/p>\n<p>To make this concrete, we ran a labored instance: take the bottom mannequin and sharpen it on two mid-resource European languages \u2014 Greek, and Bulgarian \u2014 then measure actually on held-out knowledge. The outcomes beneath are from that run. This part is a high-level overview and the coding instance lives within the companion GitHub repo. Once we publish an agentic SKILL.md protecting the entire course of, this weblog can be up to date accordingly.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhy fine-tune?<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>A number of conditions the place it pays off:<\/p>\n<p>Sharpening a long-tail locale. Languages with much less pretraining knowledge have probably the most to realize.<br \/>\nArea experience or specialised vocabulary Medical, authorized, monetary, or technical vocabulary the bottom mannequin hardly ever noticed.<br \/>\nAccent, dialect, and acoustics. Telephony, far-field, in-car, or a selected speaker inhabitants.<br \/>\nNew languages. Bootstrapping a locale that is not but lined.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tA Preview of the Energy of Fantastic-Tuning<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p><img decoding=\"async\" src=\"https:\/\/img.youtube.com\/vi\/kP9yaH-DT8E\/maxresdefault.jpg\" alt=\"Watch the Nemotron 3.5 ASR Fine-Tuning Walkthrough\"\/><\/p>\n<p>\ud83c\udfa5 Video Walkthrough: Watch on YouTube<\/p>\n<p>This walkthrough demonstrates multilingual streaming inference, latency\/accuracy tradeoffs, deployment choices, and the fine-tuning workflow described beneath.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe recipe at a look<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The entire workflow is 5 strikes:<\/p>\n<p>Level the coach at tarred speech knowledge for the goal languages \u2014 no per-file unpacking, streamed effectively by NeMo\/Lhotse.<br \/>\nFantastic-tune from the bottom checkpoint (init_from_nemo_model) utilizing the identical Cache-Conscious FastConformer-RNNT recipe, conditioned on every clip&#8217;s language tag.<br \/>\nConsider on a held-out set the mannequin by no means noticed \u2014 on the identical low-latency streaming setting you may deploy (e.g. att_context_size=[56,0], 80ms chunk; 0ms lookahead).<br \/>\nAdd extra knowledge the place the language is weak and retrain.<br \/>\nExport and deploy the fine-tuned checkpoint.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tStep 1 \u2014 Knowledge<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>We assembled a balanced, ~2000-hour combine throughout the 2 languages (Greek and Bulgarian) from public multilingual corpora (Granary, Widespread Voice, FLEURS), saved as tarred NeMo\/Lhotse shards. The 2 particulars that matter most:<\/p>\n<p>Each clip carries a target_lang tag \u2014 that is what drives the mannequin&#8217;s prompt-based language conditioning, so getting the tag proper (and utilizing a worth the mannequin acknowledges) is crucial.<br \/>\nMatch the bottom mannequin&#8217;s textual content type \u2014 punctuated, properly-cased transcripts, since that is what the mannequin produces.<\/p>\n<p>Held-out FLEURS check splits (which weren&#8217;t in coaching) gave us an sincere, in-the-wild benchmark per language.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tStep 2 \u2014 Practice<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>A simple full fine-tune of the streaming RNNT mannequin, pushed by a set step price range (the suitable strategy to schedule with streaming\/iterable knowledge). It runs on a single GPU for a fast move and scales cleanly to multi-GPU for a fuller run. On a small dataset like this, an epoch is minutes, not hours.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tStep 3 \u2014 Consider<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>We measured Phrase Error Fee on the held-out FLEURS check set, in streaming mode with 80ms chunk \u2014 probably the most demanding situation, with no future-audio &#8220;peeking.&#8221; The advance over the bottom mannequin is massive, particularly for the languages that started off weakest:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Language<br \/>\nBase mannequin<br \/>\nFantastic-tuned<br \/>\nRelative Enchancment in WER<\/p>\n<p>\ud83c\uddec\ud83c\uddf7 Greek<br \/>\n35<br \/>\n24<br \/>\n32%<\/p>\n<p>\ud83c\udde7\ud83c\uddec Bulgarian<br \/>\n22<br \/>\n15<br \/>\n31%<\/p>\n<\/div>\n<p>Uncooked WER (%) on held-out FLEURS check, lowest-latency streaming. Similar analysis for each the bottom and the fine-tuned fashions.<\/p>\n<p>Languages with increased error charges within the base mannequin grew to become genuinely helpful after a brief fine-tune \u2014 Bulgarian error charges greater than halved.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tStep 4 \u2014 Scale the information the place it helps<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>To check how way more knowledge goes, we then blended in ~2,000 further hours of parliamentary speech (MOSEL\/VoxPopuli) a part of the Granary Dataset, taking the coaching pool from ~290 hours to ~2,300 hours. Even partway by way of that longer run, the weakest languages improved additional (e.g. Bulgarian dropping into the high-20s), confirming the apparent lever: extra in-language knowledge retains serving to \u2014 although positive factors are uneven throughout languages and domains, so measure relatively than assume.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tStep 5 \u2014 Deploy<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The fine-tuned mannequin is identical structure as the bottom, so it drops straight into the identical serving path and also you decide your latency\/accuracy working level at inference time through att_context_size, precisely as in Half 1.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhat we realized<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Fantastic-tuning is transformative for under-resourced languages \u2014 the most important wins got here the place the bottom mannequin was weakest.<br \/>\nConsider at deployment latency, on held-out knowledge. Coaching-set scores flatter you; a separate check set at 0 ms look-ahead tells the reality.<br \/>\nGet the language tag proper. The immediate conditioning is highly effective however unforgiving of mismatched language labels.<br \/>\nDefend the opposite languages. When specializing in a multilingual mannequin, mix in a slice of the mannequin&#8217;s different languages (&#8220;replay&#8221;) and re-check them, so that you sharpen your goal locales with out eroding the remainder.<br \/>\nExtra knowledge helps, erratically. Including hours reliably moved most languages; one plateaued \u2014 a reminder that area match issues as a lot as uncooked amount.<\/p>\n<p>\ud83d\udce6 The complete walkthrough \u2014 knowledge prep scripts, coaching configs, the precise instructions, and the whole benchmark numbers \u2014 is within the companion GitHub repo. This part is the overview; the repo is the construct.<\/p>\n<p>For manufacturing serving, look out for the NIM launch later this month, offering gRPC streaming, and assist throughout NVIDIA Ampere, Hopper, Blackwell, Lovelace, Turing, Volta, and Jetson.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhat you possibly can construct with it<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>A number of of the use instances this mannequin unlocks:<\/p>\n<p>Sub-second voice brokers \u2014 ASR \u2192 LLM \u2192 TTS loops the place the speech-to-text leg is not the bottleneck.<br \/>\nDwell multilingual assembly captions \u2014 one stream, individuals in numerous languages, captions in actual time.<br \/>\nName-center analytics at international scale \u2014 one ASR backend as an alternative of a per-language vendor sprawl.<br \/>\nActual-time captioning + translation for livestreams and occasions.<br \/>\nOn-device transcription on Jetson for privacy-sensitive or disconnected environments.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tGet Began<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Able to construct multilingual speech purposes with a single streaming ASR mannequin?<\/p>\n<p>\ud83e\udd17 Attempt Nemotron 3.5 ASR: nvidia\/nemotron-3.5-asr-streaming-0.6b\ud83e\udde0 Run and fine-tune with NVIDIA NeMo: github.com\/NVIDIA-NeMo\/NeMo\ud83d\udcda Discover the coaching instance: Fantastic-Tuning Pocket book<\/p>\n<p>Whether or not you are constructing voice brokers, multilingual captioning programs, contact-center analytics, or on-device speech purposes, Nemotron 3.5 ASR supplies a single multilingual mannequin that may be deployed, personalized, and fine-tuned on your use case.<\/p>\n<p>We might like to see what you construct. Share your benchmarks, fine-tuning outcomes, and language variations on the mannequin dialogue web page:<\/p>\n<p>\ud83d\udcac Mannequin Discussions: https:\/\/huggingface.co\/nvidia\/nemotron-3.5-asr-streaming-0.6b\/discussions<\/p>\n<p>Mannequin: https:\/\/huggingface.co\/nvidia\/nemotron-3.5-asr-streaming-0.6b<\/p>\n<p>License: OpenMDW-1.1<\/p>\n<p>Runtime: NeMo 26.06+  <\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/nvidia\/fine-tuning-nemotron-35-asr\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Introducing NVIDIA Nemotron 3.5 ASR, streaming multilingual: a 600M-parameter speech-to-text mannequin that transcribes 40 language-locales from a single checkpoint, in actual time, with punctuation and capitalization inbuilt. It&#8217;s the successor of the favored Nemotron 3 ASR mannequin (English solely) which was launched on Hugging Face and as a NIM earlier this 12 months. Since its [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":70,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[52,49,51,47,50,48],"class_list":["post-68","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-accent","tag-asr","tag-domain","tag-finetune","tag-language","tag-nemotron"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent - Future News 24<\/title>\n<meta name=\"description\" content=\"A Blog post by NVIDIA on Hugging Face\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A Blog post by NVIDIA on Hugging Face\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-04T12:59:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-04T18:41:18+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent\",\"datePublished\":\"2026-06-04T12:59:00+00:00\",\"dateModified\":\"2026-06-04T18:41:18+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/\"},\"wordCount\":2154,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/nvidia\\\/fine-tuning-nemotron-35-asr.png\",\"keywords\":[\"Accent\",\"ASR\",\"Domain\",\"FineTune\",\"Language\",\"Nemotron\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/\",\"name\":\"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/nvidia\\\/fine-tuning-nemotron-35-asr.png\",\"datePublished\":\"2026-06-04T12:59:00+00:00\",\"dateModified\":\"2026-06-04T18:41:18+00:00\",\"description\":\"A Blog post by NVIDIA on Hugging Face\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#primaryimage\",\"url\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/nvidia\\\/fine-tuning-nemotron-35-asr.png\",\"contentUrl\":\"https:\\\/\\\/cdn-thumbnails.huggingface.co\\\/social-thumbnails\\\/blog\\\/nvidia\\\/fine-tuning-nemotron-35-asr.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/04\\\/fine-tuning-nemotron-35-asr\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent - Future News 24","description":"A Blog post by NVIDIA on Hugging Face","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/","og_locale":"en_US","og_type":"article","og_title":"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent - Future News 24","og_description":"A Blog post by NVIDIA on Hugging Face","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/","og_site_name":"Future News 24","article_published_time":"2026-06-04T12:59:00+00:00","article_modified_time":"2026-06-04T18:41:18+00:00","og_image":[{"url":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent","datePublished":"2026-06-04T12:59:00+00:00","dateModified":"2026-06-04T18:41:18+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/"},"wordCount":2154,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png","keywords":["Accent","ASR","Domain","FineTune","Language","Nemotron"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/","name":"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png","datePublished":"2026-06-04T12:59:00+00:00","dateModified":"2026-06-04T18:41:18+00:00","description":"A Blog post by NVIDIA on Hugging Face","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#primaryimage","url":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png","contentUrl":"https:\/\/cdn-thumbnails.huggingface.co\/social-thumbnails\/blog\/nvidia\/fine-tuning-nemotron-35-asr.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/04\/fine-tuning-nemotron-35-asr\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Find out how to Fantastic-Tune Nemotron 3.5 ASR for Your Language, Area, or Accent"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/68","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=68"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/68\/revisions"}],"predecessor-version":[{"id":69,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/68\/revisions\/69"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/70"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=68"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=68"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=68"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}