{"id":4091,"date":"2026-08-22T13:00:00","date_gmt":"2026-08-22T13:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/"},"modified":"2026-08-22T13:59:03","modified_gmt":"2026-08-22T13:59:03","slug":"why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/","title":{"rendered":"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name)"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"has-subtitle-2-font-size wp-block-paragraph\">This submit was co-authored with Max Silfverberg (Information Scientist, AI Options Lead),\u00a0Antti\u00a0Hallavo\u00a0(Lead AI Software program Engineer),\u00a0and\u00a0Pontus\u00a0Huotari\u00a0(Lead Information Scientist). We work at Alma Media,\u00a0a Finnish digital\u00a0providers,\u00a0marketplaces\u00a0and media firm.\u00a0Considered one of our focus areas is creating AI\/ML options for actual property itemizing providers, the place understanding picture content material performs\u00a0an essential position.\u00a0<\/p>\n<\/blockquote>\n<p class=\"wp-block-paragraph\"> providers deal with\u00a0tons of\u00a0of 1000&#8217;s of listings\u00a0a 12 months.\u00a0Most of these\u00a0include dozens of photographs\u00a0with no\u00a0details about\u00a0what they present.\u00a0In the meantime, search, suggestions, and a spread of inside use circumstances all profit from understanding whether or not a photograph represents a kitchen, ground plan, or backyard.<\/p>\n<p class=\"wp-block-paragraph\">Our answer is to\u00a0routinely tag photographs with room-type and content material lessons. Our\u00a0room varieties embody\u00a0LIVING ROOM, KITCHEN, and BEDROOM. We additionally tag schematic content material like\u00a0ground plans and web site plans.\u00a0Moreover, we acknowledge realtor advertising supplies, aerial photographs,\u00a0and\u00a0backyard\u00a0photographs.\u00a0Altogether, there are 23 lessons. As\u00a0Determine\u00a01 exhibits, this can be a traditional multi-label classification activity; the identical area can embody a number of room varieties without delay.\u00a0<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/08\/clay-banks-G8kRVwlZK3Q-unsplash-1-1024x683.jpg\" alt=\"\" class=\"wp-image-680670\"\/><figcaption class=\"wp-element-caption\">Determine\u00a01.\u00a0Our\u00a0system\u00a0ought to\u00a0tag\u00a0this\u00a0picture\u00a0as\u00a0LIVING ROOM\u00a0and\u00a0STAIRCASE.\u00a0The\u00a0eating\u00a0room\u00a0exhibiting\u00a0by\u00a0a\u00a0doorway\u00a0ought to\u00a0not\u00a0have an effect on\u00a0the\u00a0class.\u00a0Picture\u00a0by\u00a0Clay Banks\u00a0on\u00a0Unsplash.\u00a0<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">On\u00a0the face of it,\u00a0this\u00a0sounds\u00a0easy,\u00a0however\u00a0we have to make some tough choices. How do you have to deal with a front room picture that exhibits a bed room by a doorway? What if the picture solely exhibits 10% front room and the remaining 90% is\u00a0eating space?\u00a0The solutions rely upon the applying.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">If we have to discover all\u00a0photographs\u00a0exhibiting\u00a0kitchens, we\u00a0additionally\u00a0need to\u00a0determine\u00a0front room photographs\u00a0that\u00a0present\u00a0a\u00a0kitchen within the background.\u00a0Nonetheless, if the person particularly\u00a0asks for kitchen photographs, we solely need to present those the place the kitchen is in focus.\u00a0To assist resolve what to return,\u00a0classification\u00a0confidence is essential. However relying on the way you implement your classifier, you\u00a0would possibly\u00a0not have entry to that info.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Picture\u00a0classifiers\u00a0may be inbuilt many\u00a0methods. The trendy default method is to run photos by a third-party API which internally makes use of\u00a0a\u00a0vision-language mannequin\u00a0(VLM)\u00a0to investigate photos and generate tags in line with a\u00a0immediate.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">One other\u00a0possibility\u00a0is to coach picture classifiers on high of open-source ViT basis fashions\u00a0like Google\u00a0SigLIP\u00a0and Meta DINO, both freezing the inspiration mannequin or fine-tuning it. Every of those\u00a0designs\u00a0comes\u00a0with\u00a0its\u00a0personal benefits and trade-offs.<\/p>\n<p class=\"wp-block-paragraph\">There\u00a0already\u00a0exists\u00a0lots\u00a0of\u00a0work\u00a0evaluating\u00a0the approaches based mostly on numerical efficiency\u00a0[1].\u00a0This weblog submit goes additional; we ask the generally ignored query: How do you have to construct picture classifiers in a enterprise context?\u00a0<\/p>\n<h2 class=\"wp-block-heading\">Three\u00a0questions\u00a0earlier than\u00a0you\u00a0practice\u00a0something<\/h2>\n<p class=\"wp-block-paragraph\">We\u00a0constructed\u00a0our\u00a0proprietary\u00a0classifiers\u00a0by fine-tuning\u00a0google\/siglip-base-patch16-224. The query is:\u00a0why\u00a0do\u00a0this?\u00a0Examine\u00a0Determine 2 for the TL;DR.\u00a0Learn on for the total\u00a0story.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/08\/classification-flowchart-896x1024.png\" alt=\"\" class=\"wp-image-680682\"\/><figcaption class=\"wp-element-caption\">Determine\u00a02.\u00a0Ought to\u00a0you\u00a0immediate an API or practice your personal classifier both with or with out fine-tuning?\u00a0Picture by\u00a0writer.\u00a0<\/figcaption><\/figure>\n<h3 class=\"wp-block-heading\">Query 1:\u00a0Immediate\u00a0an\u00a0API or\u00a0practice your personal mannequin?<\/h3>\n<p class=\"wp-block-paragraph\">The selection to categorise by prompting by an exterior API or construct your personal\u00a0classifier\u00a0closely relies on your use case.\u00a0First, you have to think about whether or not your classification activity may even be prompted. It\u2019s simple to immediate automobile and\u00a0kitchen equipment classifiers\u00a0however how about click-through\u00a0charge (CTR) for\u00a0YouTube\u00a0video thumbnails?\u00a0Right here we want a\u00a0trainable classifier as a result of we actually\u00a0don\u2019t\u00a0know what influences the\u00a0click on\u00a0determination.\u00a0Conversely,\u00a0if\u00a0you want\u00a0to extract structured JSON recordsdata\u00a0from photographs\u00a0representing\u00a0constructing schematics,\u00a0a\u00a0easy\u00a0classifier simply\u00a0received\u2019t\u00a0lower it.<\/p>\n<p class=\"wp-block-paragraph\">Our actual property use case\u00a0sits within the center. Many lessons like KITCHEN and BATHROOM are simply promptable whereas others, like HALLWAY, LOFT and ALCOVE are fuzzier and tougher to verbalize.<\/p>\n<p class=\"wp-block-paragraph\">When you resolve to coach\u00a0your personal classifiers\u00a0as\u00a0we did, you in fact want coaching knowledge,\u00a0in all probability\u00a0not less than just a few thousand examples per class. When launching a brand new product, that&#8217;s one thing you may not have. When you not less than have entry to plain photographs with out annotations,you possibly can launch with aprompted VLM as your first classifier. Its predictions steadily accumulate into an annotated dataset, which you&#8217;ll later use to coach a customized classifier.\u00a0This will likely require a\u00a0cleanup\u00a0go, because the dataset inherits the VLM\u2019s errors.<\/p>\n<p class=\"wp-block-paragraph\">Value is one other main query.\u00a0With\u00a0a\u00a0quantity\u00a0within the\u00a0tens of millions,\u00a0the totally different classification approaches end in dramatically divergent value\u00a0profiles. Utilizing Google\u2019s\u00a0Agent Platform\u00a0and the gemini-3.5-flash mannequin,\u00a0the\u00a0July 2026\u00a0worth\u00a0is\u00a0roughly\u00a0$1.50\u00a0per 1,000\u00a0photos (at\u00a01K\u00a0decision),\u00a0so classifying one million photographs prices\u00a0roughly $1,500.<\/p>\n<p class=\"wp-block-paragraph\">Utilizing our personal classifier on\u00a0a\u00a0devoted\u00a0AWS EC2\u00a0g4dn.xlarge occasion with a T4 GPU, we will\u00a0classify\u00a0not less than\u00a0400 photos\u00a0per second. At a\u00a0July 2026\u00a0on-demand\u00a0hourly\u00a0charge\u00a0of\u00a0$0.53,\u00a0classifying one million inputs\u00a0comes\u00a0out\u00a0to\u00a0$0.37\u00a0or\u00a0roughly\u00a01\/4000th\u00a0of\u00a0the\u00a0value for\u00a0the\u00a0API\u00a0answer\u00a0(inference compute\u00a0solely).<\/p>\n<p class=\"wp-block-paragraph\">Nonetheless,\u00a0if\u00a0you classify just a few hundred photographs a day, from the price\u00a0perspective\u00a0it actually doesn\u2019t matter the way you do it. Prices change into a problem solely at scale.<\/p>\n<p class=\"wp-block-paragraph\">Along with labels, classification confidence is usually helpful. As talked about above, if we provide kitchen photographs to the person, we should always in all probability go together with assured matches. It&#8217;s, nonetheless, tough to derive dependable confidence estimates from a VLM; verbalized confidence estimates are identified to be poorly calibrated [2] and token log-likelihoods from an API often don\u2019t symbolize the class-probabilities you&#8217;re really considering.<\/p>\n<p class=\"wp-block-paragraph\">When you use an API, you would possibly due to this fact must depend on granular tags like PROBABLE\/POSSIBLE\/UNLIKELY [3], and there&#8217;s no assure that these might be dependable both. When you as an alternative practice your personal classifier, you get usable per-class scores which may be calibrated when wanted. Desk 1 summarizes how the 2 approaches examine.<\/p>\n<figure class=\"wp-block-table\">Prompted VLM (API)Customized classifierTraining knowledge None neededA few 1000 examples per class Setup effortWrite a promptannotate, practice, deployCost per 1M photographs~$1,500~$0.37 (on GPU)Per-class scoresUnreliable \/ not exposedExplicit, thresholdable, calibratable Fuzzy classesHard to verbalize in promptLearnable from examplesChanging the taskEdit the promptRetrain the mannequin<figcaption class=\"wp-element-caption\">Desk\u00a01.\u00a0Comparability between VLM and customized classifier.\u00a0<\/figcaption><\/figure>\n<h3 class=\"wp-block-heading\">Query 2:\u00a0Which basis mannequin\u00a0to\u00a0use?<\/h3>\n<p class=\"wp-block-paragraph\">When you resolve to coach your personal classifier, the one affordable alternative for many is to start out with a pretrained open-source imaginative and prescient mannequin, sometimes a imaginative and prescient transformer. For enterprise use, first test that the mannequin\u2019s license permits business use.<\/p>\n<p class=\"wp-block-paragraph\">Past that, your\u00a0enterprise objective\u00a0ought to drive the selection, as a result of totally different pretraining methods\u00a0produce totally different representations:<\/p>\n<p>SigLIP [4] (Google) is educated on captioned photos, so it attends to caption-worthy issues: canine, vehicles, individuals. Its representations are extremely object-oriented; background and digital camera angle obtain far much less emphasis.<\/p>\n<p>DINO [5, 6] (Meta) is self-supervised with patch-level goals: each area of the picture contributes to the loss, not simply the caption-worthy objects. That makes it a robust candidate when background or format issues [7]. We put this to the take a look at under.<\/p>\n<p>RADIO \/ AM-RADIO [8] (NVIDIA) agglomerate representations from a number of ViT basis fashions by distillation.<\/p>\n<p>I-JEPA[9] Meta) is self-supervised like DINO however based mostly on masked prediction.<\/p>\n<h3 class=\"wp-block-heading\">Query 3: To fine-tune or to not fine-tune?<\/h3>\n<p class=\"wp-block-paragraph\">The only strategy to begin is coaching a linear classifier on high of frozen ViT representations. There are two main benefits: it&#8217;s conceptually easy and lightning quick. You&#8217;ll be able to practice on a laptop computer in a matter of minutes utilizing a 100k-instance coaching set. Sometimes, this results in very affordable efficiency.<\/p>\n<p class=\"wp-block-paragraph\">When you resolve to fine-tune, one of the best observe is to make use of low-rank adapters (LoRA), which freeze the precise ViT spine and inject just a few skinny trainable parameter layers into the mannequin [10]. After coaching, these may be merged with the unique mannequin to keep away from prices at inference time. LoRA retains coaching tractable even on a modest GPU setup, whereas delivering almost the identical efficiency acquire as full fine-tuning.<\/p>\n<p class=\"wp-block-paragraph\">Since a shallow linear classifier normally performs effectively, fine-tuning may end up in modest beneficial properties by way of uncooked F1 rating. Nonetheless, under-labeling generally is a actual downside whenever you freeze your basis mannequin as we see under.<\/p>\n<h2 class=\"wp-block-heading\">Easy has a price ticket<\/h2>\n<p class=\"wp-block-paragraph\">The key downside with the frozen mannequin is low classification confidence. At a normal 0.5 working threshold, a whopping 35% of photographs obtain no labels from the mannequin. Tuning down the edge helps, nevertheless it comes at the price of decrease precision. Per-class thresholds would possibly assist, however downstream purposes want scores that imply the identical factor throughout all 23 lessons, and class-specific thresholds would drift with each retraining.<\/p>\n<p class=\"wp-block-paragraph\">In observe, we settled on a compromise of 0.2, which offers affordable protection and precision. Determine 3 illustrates what this appears like for a single picture: at t = 0.5 nothing clears the bar, whereas at t = 0.2 the 2 right labels come by.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/08\/frozen-model-underlabeled-1024x603.png\" alt=\"\" class=\"wp-image-680684\"\/><figcaption class=\"wp-element-caption\">Determine\u00a03.\u00a0Frozen-model\u00a0confidences\u00a0for a single picture\u00a0(illustrative). At the usual threshold (t = 0.5) the picture receives no labels; reducing it to t = 0.2 recovers\u00a0LIVING ROOM\u00a0and\u00a0DINING AREA, however at the price of decrease total classification precision.\u00a0Picture by\u00a0writer.<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">A secondary downside is poor classification on just a few frequent lessons like GARDEN and HALLWAY. Edge circumstances additionally trigger issues: when a eating set is seen in a front room picture, we want to label it each LIVING ROOM and DINING AREA. Nonetheless, when the eating set is seen solely by a doorway, we don\u2019t need the DINING AREA label.<\/p>\n<p class=\"wp-block-paragraph\">These\u00a0issues may be\u00a0addressed\u00a0by\u00a0LoRA\u00a0fine-tuning.<\/p>\n<h2 class=\"wp-block-heading\">Placing\u00a0it to\u00a0the\u00a0take a look at<\/h2>\n<p class=\"wp-block-paragraph\">We determined to coach our personal classifier and in contrast the 2 customized approaches outlined above: a frozen basis mannequin mixed with a shallow linear classifier, and fine-tuning with LoRA. In each circumstances, we added 23 unbiased classification heads on high of the inspiration mannequin, one per class.<\/p>\n<p class=\"wp-block-paragraph\">The enter is a picture vector generated by SigLIP. We moreover experiment with DINOv2 as a frozen baseline to see how caption-training compares to self-supervised coaching. LoRA fine-tuning is finished solely on SigLIP. We used the unique SigLIP mannequin moderately than SigLIP 2 in these experiments; since we examine a frozen setup in opposition to fine-tuning on the identical spine, the conclusions don\u2019t hinge on the mannequin technology.<\/p>\n<p class=\"wp-block-paragraph\">For analysis, we use micro averaged F1 rating. This emphasizes efficiency on widespread lessons like KITCHEN and LIVING ROOM, that are most central for our use circumstances.<\/p>\n<p class=\"wp-block-paragraph\">Moreover, we consider protection on the take a look at set: how most of the photographs get not less than one label? Whereas there&#8217;s a pure residual of inputs that don\u2019t fall into any of the 23 lessons, we need to discover all of the photographs that may be labeled.<\/p>\n<h3 class=\"wp-block-heading\">Coaching<\/h3>\n<p class=\"wp-block-paragraph\">We practice our classifiers on our personal proprietary set of 40k manually annotated photographs, the place every enter will get 1-3 class labels. Our validation knowledge has 1.9k examples; we break up this into 100 growth and 1.8k take a look at examples. Coaching, growth and take a look at photographs come from distinct listings, so photographs of the identical property by no means seem in multiple break up.<\/p>\n<p class=\"wp-block-paragraph\">For each our frozen baselines, we educated 23 separate sklearn LogisticRegression fashions.<\/p>\n<p class=\"wp-block-paragraph\">We educated LoRA utilizing the PEFT library. Following widespread observe [10], we wrapped the SigLIP ViT self-attention question and worth layers in LoRA adapters, leaving the MLP layers untouched, and used BCE loss on high of 23 unbiased logistic classification heads. This meant coaching solely about 0.6% of the mannequin\u2019s parameters, roughly a 99% discount in comparison with full fine-tuning. Additionally it is why the entire sweep suits on a single T4.<\/p>\n<p class=\"wp-block-paragraph\">We did a random 40-trial hyperparameter sweep [11] over the configurations in Desk 2, fixing all different hyperparameters to plain values.<\/p>\n<figure class=\"wp-block-table\">HyperparameterRangeDistributionlr1e-5 -&gt; 1e-3log-uniformbatch_size{16, 32, 64}uniform categoricallora_r{8, 16, 32}uniform categorical (lora_alpha locked to lora_r)<figcaption class=\"wp-element-caption\">Desk\u00a02. Hyperparameter sweep for\u00a0LoRA\u00a0coaching.\u00a0<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">For quick and numerically safer coaching, we used combined precision with fp16 autocast and loss scaling [12]. We educated for 20 epochs and picked the mannequin that delivers one of the best F1 rating on the event set.<\/p>\n<p class=\"wp-block-paragraph\">All coaching is finished on an AWS EC2 g4dn.xlarge occasion with a single NVIDIA T4 having 16\u00a0GB VRAM.<\/p>\n<h3 class=\"wp-block-heading\">Analysis<\/h3>\n<p class=\"wp-block-paragraph\">By way of plain micro averaged F1, variations are modest. At 82.6% F1, the fine-tuned mannequin beats each frozen SigLIP\u2019s 78.4% F1 and frozen DINOv2\u2019s 78.3% F1, however the distinction is simply round 4 factors. The frozen SigLIP and DINOv2 classifiers ship basically an identical efficiency. Frozen fashions are reported at their finest dev-set thresholds (0.2 for SigLIP, 0.35 for DINOv2); the fine-tuned mannequin at its default threshold of 0.5, which marginally understates its finest achievable F1 (83.1%). Desk 3 exhibits the total outcomes.<\/p>\n<figure class=\"wp-block-table\">MetricFrozen SigLIP (t = 0.2)Frozen DINOv2 (t = 0.35)LoRA SigLIP (t = 0.5)Micro F178.478.382.6Micro precision85.185.186.2Micro recall72.872.479.3Unlabeled photos9.4percent10.8percent3.4%<figcaption class=\"wp-element-caption\">Desk\u00a03. Numerical outcomes.\u00a0<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">The rise in F1 rating is principally as a consequence of recall, which improves by roughly 7 factors from 72.8% (SigLIP) and 72.4% (DINOv2) to 79.3%. At 85.1%, the frozen fashions\u2019 precision is already very excessive, and it solely improves by about 1 level.<\/p>\n<p class=\"wp-block-paragraph\">As Determine 4 exhibits, these outcomes will not be an artifact of the working threshold; the fine-tuned classifier outperforms the frozen SigLIP classifier at each working threshold, exhibiting that fine-tuning doesn&#8217;t merely push confidence up however genuinely improves classification efficiency. With solely 100 growth examples, we deal with the chosen thresholds and stopping epoch as coarse decisions moderately than extremely tuned optima.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/08\/pr_tradeoff-1024x864.png\" alt=\"\" class=\"wp-image-680692\"\/><figcaption class=\"wp-element-caption\">Determine\u00a04. Precision\u2013recall curves.\u00a0Markers present\u00a0fashions\u2019\u00a0working factors:\u00a0t = 0.5 (SigLIP\u00a0LoRA),\u00a0t = 0.2 (SigLIP\u00a0frozen) and t = 0.35\u00a0(DINOv2 frozen).\u00a0Axes are cropped under 0.5 to give attention to the area\u00a0the place\u00a0fashions\u00a0may\u00a0moderately\u00a0be\u00a0deployed. The fine-tuned mannequin outperforms the frozen ones\u00a0in any respect working thresholds.\u00a0Picture by\u00a0writer.<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">The modest beneficial properties in micro averaged F1 conceal substantial enhancements for particular person lessons, particularly for GARDEN (help in take a look at set: 237) with a formidable 26-point rise in comparison with frozen SigLIP, and DINING AREA (help in take a look at set: 150) with a good 15-point enchancment.<\/p>\n<p class=\"wp-block-paragraph\">The GARDEN class is a very fascinating instance, as a result of it&#8217;s sometimes all background, one thing that SigLIP doesn&#8217;t do effectively off the shelf, as mentioned above. In such circumstances, fine-tuning can ship dramatic enhancements.<\/p>\n<p class=\"wp-block-paragraph\">Nonetheless, after we have a look at efficiency for the GARDEN class utilizing the frozen DINOv2 mannequin, a unique sample emerges: frozen DINOv2 F1 rating is 58%, a 15-point enchancment over the frozen SigLIP mannequin. Simply by selecting a extra appropriate basis mannequin, we&#8217;ve got gained greater than half of the efficiency hole in comparison with a fine-tuned SigLIP mannequin. DINING AREA additionally exhibits an enchancment of 5 factors F1 rating.<\/p>\n<p class=\"wp-block-paragraph\">On the similar time, DINOv2 underperforms in comparison with SigLIP on many lessons the place semantic understanding of the picture appears extra essential: it by no means predicts MARKETING (help in take a look at set: 24), and KITCHEN (help in take a look at set: 272) slips 7 factors. Curiously, efficiency additionally degrades on HALLWAY (help in take a look at set: 86), a basically architectural class the place we might have anticipated DINOv2 to excel.<\/p>\n<p class=\"wp-block-paragraph\">The one class the place fine-tuning degrades efficiency is KITCHEN: an F1 drop of 5 factors. We think about this minor, however that is naturally case dependent.<\/p>\n<p class=\"wp-block-paragraph\">The true promoting level for LoRA fine-tuning is that it largely solves under-labeling. The frozen SigLIP classifier leaves 9.4% of photographs with out labels and DINOv2 does even worse at 10.8%. SigLIP\u2019s charge is 2.4x the pure charge (3.9%) of photographs that genuinely fall into none of our 23 lessons. LoRA finally ends up at 3.4%, barely under the pure charge, that means it as an alternative very sometimes over-labels.<\/p>\n<p class=\"wp-block-paragraph\">Determine 5 exhibits that the under-labeling charge of the fine-tuned classifier stays low for affordable working thresholds. We will commerce a little bit of recall for even larger precision. In distinction, the under-labeling charge of the frozen classifier shoots towards the sky if one tries to sharpen precision by elevating the working threshold. As a classifier, it&#8217;s due to this fact far much less versatile than the fine-tuned one.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/08\/underlabeling-1024x846.png\" alt=\"\" class=\"wp-image-680693\"\/><figcaption class=\"wp-element-caption\">Determine\u00a05. Underneath-labeling charge\u00a0as a operate of working threshold.\u00a0All\u00a0fashions are marked at their\u00a0working thresholds, t = 0.5,\u00a0t =\u00a00.2\u00a0and t = 0.35, respectively. Whereas the under-labeling charge of the fine-tuned classifier\u00a0stays\u00a0modest at\u00a0working\u00a0thresholds, the frozen\u00a0fashions\u2019\u00a0charges\u00a0rise steeply. The pure under-labeling charge within the take a look at set is 3.9%.\u00a0Picture by\u00a0writer.<\/figcaption><\/figure>\n<h2 class=\"wp-block-heading\">So,\u00a0when\u00a0ought to\u00a0you\u00a0fine-tune?<\/h2>\n<p class=\"wp-block-paragraph\">There are some things value contemplating. If under-labeling is an issue for you, then fine-tuning may be value it. The issue basically disappeared in our case.<\/p>\n<p class=\"wp-block-paragraph\">For particular person lessons, we did see massive beneficial properties, particularly in recall. GARDEN and DINING AREA at the moment are acknowledged way more usually. Nonetheless, simply selecting an applicable basis mannequin (DINOv2 moderately than SigLIP) recovered greater than half of the GARDEN hole with out fine-tuning. Nonetheless, we didn&#8217;t observe degradation of precision, so beneficial properties are real albeit modest by way of uncooked F1.<\/p>\n<p class=\"wp-block-paragraph\">With a coaching set of 40k examples, the price for a full hyperparameter sweep turned out to be round $30, which is negligible. Nonetheless, if under-labeling is just not a problem, you would possibly choose to make use of a frozen spine, particularly when periodic retraining is required. Coaching 23 classification heads on a CPU takes minutes and wishes no GPU; a full LoRA sweep takes days on a devoted GPU occasion, and that value repeats each time you retrain.<\/p>\n<p class=\"wp-block-paragraph\">The bottom-effort possibility can be VLM-based classification, however at excessive volumes that turns into a big recurring value. The distinction between $1,500 and $0.37 for one million inputs provides up shortly. Nonetheless, keep in mind that coaching your personal classifier requires annotated knowledge. We use 40k manually annotated photographs. That isn&#8217;t free both.<\/p>\n<p class=\"wp-block-paragraph\">For us, the funding has already paid off. The fine-tuned classifier now runs in\u00a0manufacturing, and the labeling high quality is nice sufficient that Alma has constructed new performance on high of it. As a result of the labels are produced routinely, they&#8217;re accessible at scale for downstream purposes to construct on.<\/p>\n<h2 class=\"wp-block-heading\">Conclusion<\/h2>\n<p class=\"wp-block-paragraph\">Whichever method you select, picture labeling pays off throughout the true property itemizing service: search outcomes, suggestions, and a spread of inside use circumstances all enhance. When you\u2019re uncertain whether or not it\u2019s value it, begin small and immediate a VLM to categorise a subset of your knowledge. From there, a light-weight classification head on high of an present embedding mannequin will lower your prices, and if you happen to want extra accuracy, fine-tuning your personal mannequin is the pure closing step.<\/p>\n<h2 class=\"wp-block-heading\">References<\/h2>\n<p class=\"wp-block-paragraph\">[1]\u00a0N.\u00a0Kisel,\u00a0I.\u00a0Volkov,\u00a0Okay.\u00a0Janouskova\u00a0and\u00a0J.\u00a0Matas,\u00a0Multimodal massive language fashions as picture classifiers\u00a0(2026),\u00a0arXiv:2603.06578\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[2]\u00a0M.\u00a0Xiong,\u00a0Z.\u00a0Hu,\u00a0X.\u00a0Lu,\u00a0Y.\u00a0Li,\u00a0J.\u00a0Fu,\u00a0J.\u00a0He\u00a0and\u00a0B.\u00a0Hooi,\u00a0Can\u00a0LLMs\u00a0categorical\u00a0their\u00a0uncertainty? An\u00a0empirical\u00a0analysis\u00a0of\u00a0confidence\u00a0elicitation\u00a0in\u00a0LLMs\u00a0(2024),\u00a0Worldwide Convention on Studying\u00a0Representations\u00a0(ICLR)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[3]\u00a0S.\u00a0Lin,\u00a0J.\u00a0Hilton\u00a0and\u00a0O.\u00a0Evans,\u00a0Educating\u00a0fashions\u00a0to precise\u00a0their\u00a0uncertainty\u00a0in\u00a0phrases\u00a0(2022),\u00a0Transactions\u00a0on Machine Studying\u00a0Analysis\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[4]\u00a0X.\u00a0Zhai,\u00a0B.\u00a0Mustafa,\u00a0A.\u00a0Kolesnikov\u00a0and\u00a0L.\u00a0Beyer,\u00a0Sigmoid\u00a0loss\u00a0for\u00a0language\u00a0picture\u00a0pre-training\u00a0(2023),\u00a0IEEE\/CVF Worldwide Convention on Laptop Imaginative and prescient (ICCV)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[5]\u00a0M.\u00a0Caron,\u00a0H.\u00a0Touvron,\u00a0I.\u00a0Misra,\u00a0H.\u00a0J\u00e9gou,\u00a0J.\u00a0Mairal,\u00a0P.\u00a0Bojanowski\u00a0and\u00a0A.\u00a0Joulin,\u00a0Rising\u00a0properties in self-supervised imaginative and prescient transformers\u00a0(2021),\u00a0IEEE\/CVF Worldwide Convention on Laptop Imaginative and prescient (ICCV)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[6]\u00a0M.\u00a0Oquab,\u00a0T.\u00a0Darcet,\u00a0T.\u00a0Moutakanni,\u00a0H.\u00a0Vo,\u00a0M.\u00a0Szafraniec,\u00a0V.\u00a0Khalidov,\u00a0et al.,\u00a0DINOv2: Studying\u00a0strong\u00a0visible\u00a0options\u00a0with out\u00a0supervision\u00a0(2024),\u00a0Transactions\u00a0on Machine Studying\u00a0Analysis\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[7]\u00a0M.\u00a0El\u00a0Banani,\u00a0A.\u00a0Raj,\u00a0Okay.-Okay.\u00a0Maninis,\u00a0A.\u00a0Kar,\u00a0Y.\u00a0Li,\u00a0M.\u00a0Rubinstein,\u00a0D.\u00a0Solar,\u00a0L.\u00a0Guibas,\u00a0J.\u00a0Johnson\u00a0and\u00a0V.\u00a0Jampani,\u00a0Probing\u00a0the\u00a03D\u00a0consciousness\u00a0of\u00a0visible\u00a0basis\u00a0fashions\u00a0(2024),\u00a0IEEE\/CVF Convention on Laptop Imaginative and prescient and\u00a0Sample\u00a0Recognition\u00a0(CVPR)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[8]\u00a0M.\u00a0Ranzinger,\u00a0G.\u00a0Heinrich,\u00a0J.\u00a0Kautz\u00a0and\u00a0P.\u00a0Molchanov,\u00a0AM-RADIO:\u00a0Agglomerative\u00a0imaginative and prescient\u00a0basis\u00a0mannequin\u00a0cut back\u00a0all\u00a0domains\u00a0into\u00a0one\u00a0(2024),\u00a0IEEE\/CVF Convention on Laptop Imaginative and prescient and\u00a0Sample\u00a0Recognition\u00a0(CVPR)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[9]\u00a0M.\u00a0Assran,\u00a0Q.\u00a0Duval,\u00a0I.\u00a0Misra,\u00a0P.\u00a0Bojanowski,\u00a0P.\u00a0Vincent,\u00a0M.\u00a0Rabbat,\u00a0Y.\u00a0LeCun\u00a0and\u00a0N.\u00a0Ballas,\u00a0Self-supervised studying from photos with a joint-embedding predictive structure\u00a0(2023),\u00a0IEEE\/CVF Convention on Laptop Imaginative and prescient and Sample Recognition (CVPR)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[10]\u00a0E. J.\u00a0Hu,\u00a0Y.\u00a0Shen,\u00a0P.\u00a0Wallis,\u00a0Z.\u00a0Allen-Zhu,\u00a0Y.\u00a0Li,\u00a0S.\u00a0Wang,\u00a0L.\u00a0Wang\u00a0and\u00a0W.\u00a0Chen,\u00a0LoRA: Low-rank adaptation of enormous language fashions\u00a0(2022),\u00a0Worldwide Convention on Studying Representations (ICLR)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[11]\u00a0J.\u00a0Bergstra\u00a0and\u00a0Y.\u00a0Bengio,\u00a0Random seek for hyper-parameter optimization\u00a0(2012),\u00a0Journal of Machine Studying Analysis, 13\u00a0<\/p>\n<p class=\"wp-block-paragraph\">[12]\u00a0P.\u00a0Micikevicius,\u00a0S.\u00a0Narang,\u00a0J.\u00a0Alben,\u00a0G.\u00a0Diamos,\u00a0E.\u00a0Elsen,\u00a0D.\u00a0Garcia,\u00a0B.\u00a0Ginsburg,\u00a0M.\u00a0Houston,\u00a0O.\u00a0Kuchaiev,\u00a0G.\u00a0Venkatesh\u00a0and\u00a0H.\u00a0Wu,\u00a0Combined precision coaching\u00a0(2018),\u00a0Worldwide Convention on Studying Representations (ICLR)<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/towardsdatascience.com\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>This submit was co-authored with Max Silfverberg (Information Scientist, AI Options Lead),\u00a0Antti\u00a0Hallavo\u00a0(Lead AI Software program Engineer),\u00a0and\u00a0Pontus\u00a0Huotari\u00a0(Lead Information Scientist). We work at Alma Media,\u00a0a Finnish digital\u00a0providers,\u00a0marketplaces\u00a0and media firm.\u00a0Considered one of our focus areas is creating AI\/ML options for actual property itemizing providers, the place understanding picture content material performs\u00a0an essential position.\u00a0 providers deal with\u00a0tons of\u00a0of 1000&#8217;s [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4093,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[7],"tags":[4353,776,4352],"class_list":["post-4091","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-science-mlops","tag-call","tag-finetuned","tag-siglip"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name) - Future News 24<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name) - Future News 24\" \/>\n<meta property=\"og:description\" content=\"This submit was co-authored with Max Silfverberg (Information Scientist, AI Options Lead),\u00a0Antti\u00a0Hallavo\u00a0(Lead AI Software program Engineer),\u00a0and\u00a0Pontus\u00a0Huotari\u00a0(Lead Information Scientist). We work at Alma Media,\u00a0a Finnish digital\u00a0providers,\u00a0marketplaces\u00a0and media firm.\u00a0Considered one of our focus areas is creating AI\/ML options for actual property itemizing providers, the place understanding picture content material performs\u00a0an essential position.\u00a0 providers deal with\u00a0tons of\u00a0of 1000&#8217;s [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-22T13:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-22T13:59:03+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"18 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name)\",\"datePublished\":\"2026-08-22T13:00:00+00:00\",\"dateModified\":\"2026-08-22T13:59:03+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/\"},\"wordCount\":3531,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg\",\"keywords\":[\"Call\",\"finetuned\",\"SigLip\"],\"articleSection\":[\"Data Science &amp; MLOps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/\",\"name\":\"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name) - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg\",\"datePublished\":\"2026-08-22T13:00:00+00:00\",\"dateModified\":\"2026-08-22T13:59:03+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#primaryimage\",\"url\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg\",\"contentUrl\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/22\\\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name)\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name) - Future News 24","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/","og_locale":"en_US","og_type":"article","og_title":"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name) - Future News 24","og_description":"This submit was co-authored with Max Silfverberg (Information Scientist, AI Options Lead),\u00a0Antti\u00a0Hallavo\u00a0(Lead AI Software program Engineer),\u00a0and\u00a0Pontus\u00a0Huotari\u00a0(Lead Information Scientist). We work at Alma Media,\u00a0a Finnish digital\u00a0providers,\u00a0marketplaces\u00a0and media firm.\u00a0Considered one of our focus areas is creating AI\/ML options for actual property itemizing providers, the place understanding picture content material performs\u00a0an essential position.\u00a0 providers deal with\u00a0tons of\u00a0of 1000&#8217;s [&hellip;]","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/","og_site_name":"Future News 24","article_published_time":"2026-08-22T13:00:00+00:00","article_modified_time":"2026-08-22T13:59:03+00:00","og_image":[{"url":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg","twitter_misc":{"Written by":"Future News 24","Est. reading time":"18 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name)","datePublished":"2026-08-22T13:00:00+00:00","dateModified":"2026-08-22T13:59:03+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/"},"wordCount":3531,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#primaryimage"},"thumbnailUrl":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg","keywords":["Call","finetuned","SigLip"],"articleSection":["Data Science &amp; MLOps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/","name":"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name) - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#primaryimage"},"thumbnailUrl":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg","datePublished":"2026-08-22T13:00:00+00:00","dateModified":"2026-08-22T13:59:03+00:00","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#primaryimage","url":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg","contentUrl":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/08\/clay-banks-EskHgf31GUU-unsplash-1-scaled-1.jpg"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/22\/why-we-fine-tuned-siglip-and-why-thats-not-always-the-right-call\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Why We Nice-Tuned SigLip (And Why That\u2019s Not All the time the Proper Name)"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4091","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=4091"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4091\/revisions"}],"predecessor-version":[{"id":4092,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4091\/revisions\/4092"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/4093"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=4091"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=4091"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=4091"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}