{"id":2549,"date":"2026-07-17T00:00:00","date_gmt":"2026-07-17T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/"},"modified":"2026-07-19T04:59:04","modified_gmt":"2026-07-19T04:59:04","slug":"visual-concept-inference","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/","title":{"rendered":"Present Me Examples: Inferring Visible Ideas from Picture Units"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p>Imaginative and prescient-language fashions (VLMs) can observe complicated textual directions, but they battle to purpose from purely visible context. Specifically, present fashions fail to deduce shared ideas from units of instance photos and apply them to new inputs. We introduce Visible Idea Inference from Units (VICIS), a activity that evaluates this functionality. Given a small context set of photos sharing an idea and a question picture, the mannequin should generate new photos that protect the context-defined idea whereas remaining in step with the question. We present that state-of-the-art VLMs carry out poorly on this activity, usually ignoring the visible context or defaulting to biased generations. To handle this hole, we suggest a coaching framework and structure that be taught to deduce visible ideas from picture units and extract concept-specific embeddings from queries. Experiments on artificial information and large-scale ImageNet\/WordNet information present that our mannequin generates extra correct and various outputs and generalizes to unseen ideas and modalities resembling sketches.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/machinelearning.apple.com\/research\/visual-concept-inference\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Imaginative and prescient-language fashions (VLMs) can observe complicated textual directions, but they battle to purpose from purely visible context. Specifically, present fashions fail to deduce shared ideas from units of instance photos and apply them to new inputs. We introduce Visible Idea Inference from Units (VICIS), a activity that evaluates this functionality. Given a small [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2551,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[1281,1056,1842,3045,1948,1298,2948],"class_list":["post-2549","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-concepts","tag-examples","tag-image","tag-inferring","tag-sets","tag-show","tag-visual"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Present Me Examples: Inferring Visible Ideas from Picture Units - Future News 24<\/title>\n<meta name=\"description\" content=\"Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In\u2026\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Present Me Examples: Inferring Visible Ideas from Picture Units - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-17T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-19T04:59:04+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Present Me Examples: Inferring Visible Ideas from Picture Units\",\"datePublished\":\"2026-07-17T00:00:00+00:00\",\"dateModified\":\"2026-07-19T04:59:04+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/\"},\"wordCount\":173,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"keywords\":[\"Concepts\",\"Examples\",\"image\",\"Inferring\",\"Sets\",\"Show\",\"Visual\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/\",\"name\":\"Present Me Examples: Inferring Visible Ideas from Picture Units - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"datePublished\":\"2026-07-17T00:00:00+00:00\",\"dateModified\":\"2026-07-19T04:59:04+00:00\",\"description\":\"Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In\u2026\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"contentUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/17\\\/visual-concept-inference\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Present Me Examples: Inferring Visible Ideas from Picture Units\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Present Me Examples: Inferring Visible Ideas from Picture Units - Future News 24","description":"Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In\u2026","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/","og_locale":"en_US","og_type":"article","og_title":"Present Me Examples: Inferring Visible Ideas from Picture Units - Future News 24","og_description":"Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In\u2026","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/","og_site_name":"Future News 24","article_published_time":"2026-07-17T00:00:00+00:00","article_modified_time":"2026-07-19T04:59:04+00:00","og_image":[{"url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Present Me Examples: Inferring Visible Ideas from Picture Units","datePublished":"2026-07-17T00:00:00+00:00","dateModified":"2026-07-19T04:59:04+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/"},"wordCount":173,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","keywords":["Concepts","Examples","image","Inferring","Sets","Show","Visual"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/","name":"Present Me Examples: Inferring Visible Ideas from Picture Units - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","datePublished":"2026-07-17T00:00:00+00:00","dateModified":"2026-07-19T04:59:04+00:00","description":"Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In\u2026","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#primaryimage","url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","contentUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/17\/visual-concept-inference\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Present Me Examples: Inferring Visible Ideas from Picture Units"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2549","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2549"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2549\/revisions"}],"predecessor-version":[{"id":2550,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2549\/revisions\/2550"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2551"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2549"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2549"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2549"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}