{"id":3899,"date":"2026-08-18T07:51:00","date_gmt":"2026-08-18T07:51:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/"},"modified":"2026-08-18T09:59:13","modified_gmt":"2026-08-18T09:59:13","slug":"vram-overcommit","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/","title":{"rendered":"VRAM Administration Half 2: Past the Limits of Bodily VRAM"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div id=\"\">\n<p>  August 17, 2026<\/p>\n<p>Earlier this yr, I blogged about work I did to enhance VRAM administration<br \/>\nfor video games. Now, after many months of floating round in mailing lists, the kernel patches<br \/>\nare lastly merged upstream and queued for Linux 7.3! Hooray!<\/p>\n<p>To rejoice, let\u2019s look a bit deeper at one sentence I wrote in my earlier publish:<\/p>\n<blockquote>\n<p>[Games] ought to carry out way more steady &#8211; so long as the sport itself doesn\u2019t<br \/>\n  use extra VRAM than you even have.<\/p>\n<\/blockquote>\n<p>So, one might ask: What in the event that they do, actually, use extra VRAM than you even have?<\/p>\n<p>Typical expectations for this appear to be that when this occurs you\u2019re just about screwed.<br \/>\nVideo games will begin crashing left and proper, efficiency plummets to unplayable ranges,<br \/>\n gaming expertise turns into not possible.<\/p>\n<p>However is that basically simply an unavoidable truth of life? What actually makes operating out of<br \/>\nVRAM suck so arduous? And, most significantly: How can we make it suck as little as potential?<\/p>\n<p>In idea, operating out of VRAM ought to solely be a efficiency situation, not a stability one.<br \/>\nHelp for overcommitting VRAM has existed for so long as GPU drivers have: If the driving force<br \/>\novercommits VRAM, you might be usually allowed to request as a lot VRAM as you\u2019d like,<br \/>\nand also you\u2019ll get as a lot because the kernel driver decides it will possibly match into the bodily reminiscence that exists on GPU.<\/p>\n<p>On the efficiency facet, the big-picture motive for unhealthy efficiency whenever you run out of VRAM is pretty easy.<br \/>\nAs quickly as the sport requests extra VRAM than is bodily current, a number of the sport\u2019s reminiscence will<br \/>\nshould be moved\/evicted to CPU RAM as a substitute. For the GPU, accessing CPU RAM is far slower than VRAM:<br \/>\nNot solely is CPU RAM slower than a devoted GPU\u2019s VRAM on the whole, all reminiscence accesses additionally<br \/>\nshould go over the PCI bus. The PCI bus provides latency and is often additionally the limiting consider bandwidth when<br \/>\nfetching from CPU reminiscence.<\/p>\n<p>Attributable to PCI velocity limitations, there are some actually unavoidable efficiency constraints when overcommitting VRAM.<br \/>\nAssuming the GPU is attached through a PCIe 4.0&#215;16 connection, you get rather less than 32GiB\/s of bandwidth. Every<br \/>\nmillisecond, that PCIe bus can switch ~32.2MiB of information. For a minimal framerate of 30 frames per second (33.3ms<br \/>\nper body), absolutely the most quantity of information the GPU is ready to entry is ~1,075.5MiB, a tiny bit over 1GiB of<br \/>\ninformation. In different phrases, if a lot reminiscence will get evicted that the GPU must fetch greater than 1GiB from evicted reminiscence<br \/>\nin a single single body, it&#8217;s merely not possible to nonetheless hit 30 FPS.<\/p>\n<h2 id=\"not-all-memory-is-equal\">Not all reminiscence is equal<\/h2>\n<p>On the identical time, simply studying a little bit little bit of CPU reminiscence on the GPU shouldn&#8217;t be instantly a dying sentence for efficiency.<br \/>\nIn actual fact, GPU drivers typically determine to let issues like command buffer information and associated allocations stay in CPU<br \/>\nRAM even when there\u2019s loads of VRAM obtainable! Each time the GPU executes these instructions, it has to entry CPU reminiscence, and but<br \/>\nin these instances all the things runs fully wonderful. So what makes these accesses totally different &#8211; why are they wonderful and but<br \/>\noperating out of VRAM appears catastrophic?<\/p>\n<p>One factor that influences the calculus considerably is caching. Because the entry latency in case of a cache hit is identical<br \/>\nno matter whether or not the cached reminiscence lives on CPU or GPU, the excessive preliminary price of fetching over the PCI bus will be<br \/>\namortized by cache hits (to some extent). We will estimate latency variations between fetching CPU RAM and VRAM by writing microbenchmarks<br \/>\nthat measure entry latency for various buffer sizes<br \/>\n(utilizing an adversarial entry sample to reduce cache hitrates so far as potential). The consequence you get might<br \/>\nlook one thing like this (captured on RDNA3):<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png\" alt=\"Cache microbenchmark for an RDNA3 GPU.\"\/><\/p>\n<p>As anticipated, if the buffer matches into L2 (or any higher-level cache), entry latencies are precisely the identical for<br \/>\nreminiscence backed by CPU RAM and reminiscence backed by VRAM, as a result of the information will get fetched instantly from cache in both case.<br \/>\nAt a measurement of 6MB (the L2 cache measurement on RDNA3), CPU reminiscence latencies go as much as about 2400 cycles per entry,<br \/>\nwhereas system reminiscence latencies keep inside the identical tough ballpark.<br \/>\nObserve that VRAM accesses additionally undergo the Infinity Cache, however CPU reminiscence accesses don&#8217;t (they hit PCIe instantly on an L2 miss).<br \/>\nI believe it&#8217;s because the Infinity Cache sits instantly on prime of VRAM, so any entry that doesn\u2019t hit VRAM additionally doesn\u2019t attain<br \/>\nthe Infinity Cache.<\/p>\n<p>Clearly, reminiscence doesn\u2019t begin off with being cached wherever, so the primary entry will nonetheless have significantly greater<br \/>\nlatency. Additionally, shedding the Infinity Cache undoubtedly hurts as effectively: PCIe fetches appear to have someplace round 7.3x as a lot<br \/>\nlatency than an Infinity Cache hit, and round 4.6x as a lot latency as a fetch from VRAM.<br \/>\nThis elevated latency wants actually excessive cache hitrates to totally amortize the price of going over PCIe.<br \/>\nWhich means there&#8217;s solely a small set of use instances the place utilizing CPU reminiscence has such minuscule slowdowns<br \/>\nthat you simply\u2019d actively determine to make use of it in favor of VRAM when you might have the selection. Whenever you\u2019re evicting reminiscence from VRAM, there<br \/>\nwill nearly unavoidably be not less than some extent of slower efficiency.<\/p>\n<p>Nonetheless, although slowdown is unavoidable, there&#8217;s going to be reminiscence the place eviction issues extra and reminiscence<br \/>\nthe place eviction has a lesser impact on total perf. Reminiscence that&#8217;s accessed in very cache-friendly methods shouldn&#8217;t be<br \/>\naffected by the slowdown of CPU RAM as a lot. If the entry patterns aren\u2019t cache-friendly however the reminiscence isn\u2019t accessed very<br \/>\nusually, issues may nonetheless be wonderful for the reason that GPU solely not often wants to truly fetch information from CPU RAM. There could be many<br \/>\nreminiscence allocations the place the GPU will solely entry a small a part of the full allocation measurement, and by no means even learn the remaining.<br \/>\nIf these allocations have been to be evicted, you may evict a number of GiBs of information, however nonetheless stay effectively beneath the 1GiB arduous restrict of information<br \/>\nthat&#8217;s truly accessed per body.<\/p>\n<p>All of those variables make it surprisingly arduous to foretell how efficiency truly pans out in apply when reminiscence is being<br \/>\nevicted. However in brief: Relying on how a lot the evicted reminiscence will get accessed and the way effectively these accesses cache, you may simply have the ability to run out<br \/>\nof VRAM with out (fully) ruining efficiency!<\/p>\n<p>We\u2019ve theorycrafted ourselves all the best way in the direction of having performant VRAM overcommitment now. Nice! Let\u2019s simply boot up SteamOS, begin some sport and crank<br \/>\nup the setti-radv\/amdgpu: Not sufficient reminiscence for command submission.<\/p>\n<p>oh.<\/p>\n<p>Because it seems, operating out of VRAM in apply does carry loads of stability points with it.<\/p>\n<p>This error isn\u2019t fairly like an everyday \u201ccouldn\u2019t allocate, out of reminiscence\u201d error, although. Observe that the message particularly complains about<br \/>\ncommand submission: RADV prints this message when the kernel returns -ENOMEM when attempting to submit instructions, however merely submitting<br \/>\ninstructions doesn&#8217;t allocate any new assets! All of the command buffers have been allotted prematurely, and clearly their allocation succeeded.<br \/>\nThough all reminiscence was efficiently allotted, utilizing it in a GPU submission abruptly ends in \u201cout of reminiscence\u201d errors being thrown.<\/p>\n<p>It\u2019s time for one more kernel journey! Certainly getting the kernel to just accept the submission can\u2019t be that arduous &#8211; in spite of everything, the kernel already<br \/>\naccepted all of the allocations!<\/p>\n<h2 id=\"the-horrors-of-kernel-locking\">The horrors of kernel locking<\/h2>\n<p>One factor the amdgpu driver has to do on each submission, earlier than it will possibly direct the GPU to start out executing instructions, is to be sure that<br \/>\nall reminiscence which will doubtlessly be referenced by the GPU instructions is accessible. With extra fashionable bindless graphics APIs, it&#8217;s important to assume<br \/>\nall allotted reminiscence might in some unspecified time in the future get referenced. Subsequently, amdgpu will attempt to ensure all allotted reminiscence can be accessible.<\/p>\n<p>Every reminiscence allocation carries details about which kind of reminiscence (for our functions right here, system RAM or GPU VRAM) it may be correctly accessed<br \/>\nfrom. Most allocations will be accessed from both CPU RAM or VRAM, and amdgpu will probably be proud of the reminiscence allocation being in both of those<br \/>\nreminiscence sorts. Some allocations, nevertheless, should be positioned in VRAM and VRAM solely. If these reminiscence allocations have been evicted to system RAM as a result of<br \/>\nanother utility allotted VRAM within the meantime, amdgpu must transfer them again into VRAM. As a result of there isn&#8217;t a free VRAM obtainable in any respect,<br \/>\ntransferring the allocation again requires evicting one thing else. For some motive, that failed and the kernel<br \/>\nreported an out-of-memory situation.<\/p>\n<p>So as to clarify why evicting one thing randomly fails, we\u2019ll should take a small detour to take a look at how the kernel handles (CPU-side) locking for GPU allocations.<br \/>\nSo as to evict a reminiscence allocation, it&#8217;s important to purchase a lock related to that allocation. Nonetheless, throughout a submission, you additionally should lock<br \/>\neach allocation that\u2019s referenced in a submission, to stop another utility from transferring the allocation some other place when you\u2019re busy getting ready<br \/>\nGPU work. But when one other GPU submission is doing the identical factor concurrently, you may find yourself in a scenario like this:<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/AMDGPUDeadlock.png\" alt=\"Deadlock condition when concurrent submits evict\"\/><\/p>\n<p>If one submit desires to evict an allocation that one other submit has already locked, however that different submit additionally must lock an allocation from the primary one<br \/>\nto make progress, we now have a textbook ABBA impasse situation.<\/p>\n<p>However concern not, the kernel is aware of find out how to detect and resolve deadlocks! The main points about how impasse detection works are<br \/>\ndefined on this kernel documentation web page, however in very broad strokes, the<br \/>\nkernel associates locking operations with a \u201ctransaction\u201d (which mainly simply retains observe of which locks have been acquired). If two transactions would<br \/>\nimpasse, one of many transactions is marked as \u201cwounded\u201d, and the subsequent time it tries to amass a lock, the -EDEADLCK error is returned.<br \/>\nThis error requests the transaction to be aborted: All locks acquired throughout the transaction needs to be launched, and the transaction is restarted<br \/>\nfrom scratch. Within the context of command submission, this simply means the driving force will restart the method of going over all reminiscence allocations and making<br \/>\npositive they\u2019re accessible.<\/p>\n<p>So the place\u2019s the catch? There isn\u2019t one. This method is rock stable and works rather well.<\/p>\n<p>At the least so long as it\u2019s truly carried out in all places.<\/p>\n<p>Within the graphics subsystem, the gritty internals of the wound-abort-retry loop are abstracted utilizing a small helper library referred to as drm_exec. As an alternative of<br \/>\nhaving to manually observe which allocations are locked, and launch the locks when you run into -EDEADLCK, you merely use the drm_exec_lock_obj helper.<br \/>\nIn the event you research the locking code in TTM,<br \/>\nthe shared Linux GPU reminiscence administration layer, you&#8217;ll discover a profound lack of utilization of drm_exec.<\/p>\n<p>As an alternative, there even is a remark noting that -EDEADLCK will trigger eviction to fail. There we go, we discovered our situation! As quickly as this impasse<br \/>\nsituation is encountered due to intense reminiscence strain throughout command submission, the kernel bails out and rejects the submission as a substitute of<br \/>\nretrying.<\/p>\n<p>There already are some patchsets to hook up the drm_exec helper in TTM,<br \/>\ndespatched all the best way again in 2024, however these by no means made it in for just a few causes,<br \/>\namongst which have been some remaining bugs that hadn\u2019t been discovered. My work had been reduce out for me right here: Rebase the patchset on prime of<br \/>\nmy kernel model and work out what these remaining bugs are.<\/p>\n<p>Rebasing the patchset wasn\u2019t an excessive amount of of a trouble, and determining the bugs solely took one single week of intense struggling with video games randomly hanging 3 minutes<br \/>\ninto heavy VRAM rivalry. Not the worst!<\/p>\n<p>I attempted resending the patchset with fixes for all bugs I discovered within the<br \/>\nhopes it will get on this time, however there\u2019s going to be extra work needing to be finished with it earlier than it may be merged.<\/p>\n<p>Now that operating out of VRAM not less than received\u2019t crash your apps at random, we are able to not less than correctly crank up the settings and have a look at perf. The preliminary<br \/>\nconsequence gave me a fully wonderful efficiency graph like this:<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/MangohudVRAM.png\" alt=\"Horrible perf graph\"\/><\/p>\n<h2 id=\"hold-on-where-did-all-the-perf-go\">Maintain On The place Did All The Perf Go<\/h2>\n<p>Determining why efficiency is so rubbish requires determining what the system is definitely doing that\u2019s this gradual. For broad \u201cwhat\u2019s the kernel driver doing??\u201d<br \/>\nquestions like that, I like utilizing gpuvis. gpuvis makes use of kernel tracepoints to construct a timeline of issues that occurred<br \/>\n(together with \u201cGPU work submission began\/stopped\u201d, from which the time taken for every submission will be inferred).<\/p>\n<p>Booting up gpuvis with a hint taken whereas the system is operating out of VRAM, the timeline reveals a scenario like this:<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/gpuvisGarbage.png\" alt=\"Horrible perf with SDMA being busy most of the time\"\/><\/p>\n<p>Seems, most of that point isn\u2019t truly spent on dealing with the submission (that\u2019s the gfx_0.0.0 exercise), however as a substitute transferring round reminiscence in preparation<br \/>\nfor that submission (sdma0 exercise)!<\/p>\n<p>The explanation why there are such a lot of buffer strikes on a regular basis turns into extra apparent should you use gpuvis\u2019s occasion record, along with a filter to indicate solely captured<br \/>\ntransfer occasions for a selected buffer object (I selected one at random right here, most buffer objects have an identical sample):<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/gpuvisGarbageBO.png\" alt=\"gpuvis event filter showing buffer go ping pong ping pong ping pong\"\/><\/p>\n<p>The record reveals fairly clearly that contending processes (on this case, gamescope and the sport itself) will continually take turns evicting and transferring<br \/>\nagain the identical piece of reminiscence, time and again. That\u2019s actually unhealthy! And it\u2019s very harking back to one thing I wrote in my first blogpost:<\/p>\n<blockquote>\n<p>Typically, two competing purposes will be anticipated to roughly take turns executing GPU work &#8211; first one utility submits work, then the opposite, then the primary once more, and so forth.<br \/>\nWith that method, reminiscence would preserve being moved backwards and forwards after each single submission. One utility will get kicked out and instantly moved again in, kicking the opposite out (which strikes reminiscence again within the subsequent step). All this transferring ended up with<br \/>\nworse efficiency than if the reminiscence had by no means been moved within the first place.<\/p>\n<\/blockquote>\n<p>This described an previous situation the place overly aggressive VRAM allocation would result in ping-pong-like strikes occurring continually. However that situation had since been<br \/>\nmounted by merely not attempting to say VRAM when there isn\u2019t any free VRAM left, and the kernel solely began being considerably aggressive after I carried out<br \/>\nVRAM safety with dmem cgroups. Clearly, this will need to have reintroduced the ping-ponging someway.<\/p>\n<p>Conceptually, the design of the dmem cgroup VRAM safety ought to by no means end in ping-pong strikes, as a result of the kernel is just presupposed to evict reminiscence that<br \/>\ndoesn&#8217;t have any cgroup VRAM safety related to it. With none VRAM safety, it is best to sometimes not be allowed to evict protected VRAM.<\/p>\n<p>The one exception to this rule is reminiscence that completely has to stay in VRAM for issues to work correctly. These sorts of reminiscence allocations are all the time<br \/>\nallowed to be moved to VRAM to make sure system stability. Sometimes, nearly nothing coming from an utility is absolutely required to stay in VRAM for proper operation,<br \/>\nhowever there&#8217;s one buffer object coming from an utility that does: The buffer containing picture information to be scanned out to the show.<\/p>\n<h2 id=\"display-hardware-is-funky\">Show {hardware} is funky<\/h2>\n<p>Not solely does the show {hardware} like scanned-out photographs to be in VRAM, it additionally fully skips previous the GPU\u2019s digital reminiscence structure and<br \/>\nworks with bodily addresses solely. In consequence, scanned-out photographs additionally should be contiguous in bodily reminiscence.<\/p>\n<p>With digital reminiscence and the facility of web page tables, typical utility buffers are solely contiguous in digital reminiscence, and could also be scattered round<br \/>\nthroughout bodily reminiscence. The primary web page of a buffer at digital tackle 0x5000 could also be mapped within the web page tables to level to bodily tackle<br \/>\n0x1234000, however the second web page at digital tackle 0x6000 may level to bodily tackle 0x4321000, someplace fully totally different!<\/p>\n<p>Here&#8217;s a diagram visualizing the mapping of digital allocations to bodily ones in case the place there&#8217;s a variety of fragmentation (which usually<br \/>\nis the case whenever you\u2019re very low on VRAM):<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/AMDGPUFragmentation.png\" alt=\"Contiguous virtual memory mapping to fragmented physical memory.\" width=\"80%\"\/><\/p>\n<p>The arrows present web page desk mappings to bodily reminiscence segments for the totally different segments of the primary allocation.<br \/>\nThey\u2019re not noted for all different allocations for readability.<\/p>\n<p>In the event you\u2019re allocating show scanout information, this fragmentation shouldn&#8217;t be an possibility because the bodily reminiscence needs to be contiguous. This has very, very<br \/>\nunlucky interactions with eviction of different information particularly. Let\u2019s assume the scanout information has already been evicted, however now it\u2019s<br \/>\ntime for that information to be scanned out, so it needs to be moved again into VRAM.<\/p>\n<p>Merely evicting one buffer received\u2019t be adequate, even when that buffer is identical measurement because the show scanout information, as a result of evicting it doesn&#8217;t<br \/>\nend in sufficient contiguous bodily house to position the scanout information in! To make issues worse, the eviction algorithm doesn&#8217;t take into<br \/>\naccount bodily reminiscence constraints in any respect. It&#8217;s a very simplistic loop alongside the traces of<\/p>\n<div class=\"language-c highlighter-rouge\">\n<div class=\"highlight\"><span class=\"k\">whereas<\/span> <span class=\"p\">(<\/span><span class=\"nb\">true<\/span><span class=\"p\">)<\/span> <span class=\"p\">{<\/span><br \/>\n   <span class=\"n\">evict<\/span><span class=\"p\">(<\/span><span class=\"n\">getLeastRecentlyUsedBuffer<\/span><span class=\"p\">())<\/span><br \/>\n   <span class=\"k\">if<\/span> <span class=\"p\">(<\/span><span class=\"n\">tryAllocate<\/span><span class=\"p\">(<\/span><span class=\"n\">newBuffer<\/span><span class=\"p\">)<\/span> <span class=\"o\">==<\/span> <span class=\"n\">SUCCESS<\/span><span class=\"p\">)<\/span><br \/>\n      <span class=\"k\">break<\/span><span class=\"p\">;<\/span><br \/>\n<span class=\"p\">}<\/span>\n<\/div>\n<\/div>\n<p>Utilizing this algorithm (assuming the allocations are organized in LRU order), even should you evict the primary 3 allocations (inexperienced, blue, and pink), there received\u2019t be a big<br \/>\nsufficient house to carry the scanout buffer! Even the most important potential free house is ever so barely too small, as is seen on this up to date diagram:<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/AMDGPUFragmentationEvicted.png\" alt=\"Still no space for the scanout buffer.\" width=\"80%\"\/><\/p>\n<p>To seek out a big sufficient bodily contiguous reminiscence area in our instance, each single allocation in VRAM would find yourself being evicted! In<br \/>\nreal-world situations, I noticed as much as 4GiB of VRAM being nuked simply to create space for scanout photographs (that are ~32MiB of pixel information per picture for a R11G11B10<br \/>\npixel format). That\u2019s going to harm actual arduous! Merely the act of transferring all that information out from VRAM would already price not less than ~130ms, in keeping with<br \/>\nthe PCIe switch fee estimated earlier.<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/memes\/scanout.png\" alt=\"welp.\" height=\"300\"\/><\/p>\n<h2 id=\"throwing-heuristics-at-the-problem\">Throwing heuristics on the drawback<\/h2>\n<p>Whereas scanout is unquestionably probably the most egregious failure case right here, this situation is extra basic: There are all the time going to make certain reminiscence allocations<br \/>\nthat will probably be moved to VRAM time and again, doubtlessly kicking out some reminiscence that an utility may desire to remain in VRAM. Resisting this and<br \/>\nattempting to maneuver the evicted reminiscence again in will almost definitely backfire.<\/p>\n<p>Though dmem cgroup safety shouldn&#8217;t be a whole answer to this drawback, it does scale back the issue scope by loads. With cgroup safety, you<br \/>\ncan make certain that any random app received\u2019t attempt to kick out essential sport assets willy-nilly. Any reminiscence that does get moved again into VRAM by power<br \/>\nin all probability has  motive to be in VRAM. Subsequently, even with dmem cgroup safety, we needs to be cautious and never attempt to reclaim evicted reminiscence again<br \/>\nby power.<\/p>\n<p>With some iterative testing, I feel I\u2019ve arrived at a set of heuristics that work moderately effectively for many instances a sport would encounter within the wild<br \/>\n(not being too aggressive when stuff will get evicted by essential system allocations is one factor, nevertheless it additionally must be moderately fast at reclaiming<br \/>\nevicted reminiscence if e.g. the sport is paused and the Steam menu runs as a substitute, evicting a number of sport reminiscence, after which the sport is resumed).<\/p>\n<p>The heuristics work one thing like this:<\/p>\n<p>  When the kernel detects an utility\u2019s reminiscence is being evicted, it enters a \u201carduous throttle\u201d part for just a few milliseconds. Throughout this part, it does<br \/>\nnot attempt transferring any reminiscence for that app again into VRAM by any means (so long as all reminiscence will be correctly accessed, after all).<br \/>\n  After this era, it switches a \u201cmushy throttle\u201d part, throughout which it could reclaim free house by transferring issues again into VRAM, however doesn&#8217;t<br \/>\nattempt evicting any reminiscence that different apps have allotted. This era might final up to a couple seconds, to make further positive all the things reached a steady state.<br \/>\n  If the \u201cmushy throttle\u201d part has accomplished with none additional reminiscence being evicted once more, the system is assumed to have reached a reasonably steady state<br \/>\nand restrictions on evicting different purposes\u2019 reminiscence are eliminated.<\/p>\n<p>IME, this achieves an appropriate stability between not capturing oneself within the foot with overaggressive eviction of different apps, whereas nonetheless recovering<br \/>\nmoderately quick when a number of your reminiscence was abruptly evicted, for instance as a result of the sport was paused and the consumer browsed round on Steam as a substitute of<br \/>\ntaking part in.<\/p>\n<h2 id=\"getting-somewhere\">Getting someplace<\/h2>\n<p>With these heuristics in place, let\u2019s lastly attempt cranking up the settings for actual this time.<\/p>\n<p>I ended up going with Indiana Jones: The Nice Circle, because it conveniently exposes a setting for streaming pool sizes that you simply<br \/>\ncan mess with to switch VRAM consumption just about instantly.<\/p>\n<p>Lo and behold, even when the settings are turned as much as a considerably ridiculous level, the place the sport requests 9GiB of 8GiB VRAM (aka. a complete 1GiB of<br \/>\novercommitted sport assets dwelling in CPU reminiscence), efficiency isn\u2019t cratering into oblivion anymore! A 19.6ms per body common is what I\u2019d nonetheless<br \/>\nname completely playable.<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/IndyOvercommit.png\" alt=\"Indiana Jones: TGC overcommit to 9GiB\/8GiB\"\/><\/p>\n<p>I may bump the settings to much more ridiculous ranges and double the quantity of overcommitted reminiscence, with the sport requesting 10GiB of VRAM<br \/>\non this 8GiB system (and thus 2GiB of assets being overcommitted). Frametime variance goes up rather a lot at this level, with spikes<br \/>\nreaching above 33.3ms occurring regularly. The general common is round 29.8ms which isn\u2019t the worst, however particularly paired with the variance,<br \/>\nthis could begin being noticeable in gameplay.<\/p>\n<p>Whereas that is already an enormous step ahead, we aren\u2019t fairly there but. The expertise underneath VRAM overcommit can typically nonetheless be a bit hit-or-miss,<br \/>\nand frametimes might noticeably differ relying on which objects within the sport you\u2019re .<\/p>\n<p>Keep in mind that for truly good eviction efficiency, it issues loads how the evicted reminiscence is utilized by the GPU.<br \/>\nProper now, this isn\u2019t taken under consideration in any respect! If we have been capable of base our eviction selections extra on how effectively the appliance\u2019s accesses work<br \/>\nwith CPU reminiscence, a variety of this variance may merely disappear.<\/p>\n<p>The sophisticated factor concerning the utility\u2019s reminiscence entry patterns is that they&#8217;re solely actually identified to the appliance.<br \/>\nSubsequently, the driving force isn\u2019t actually capable of take them under consideration as-is. Ideally there could be some API the place the appliance<br \/>\ncan provide hints to the driving force about how effectively a selected reminiscence allocation is suited to being evicted.<\/p>\n<p>One thing precisely like vkSetDeviceMemoryPriorityEXT!<br \/>\nThe VK_EXT_pageable_device_local_memory extension offers exactly what we want right here, by permitting purposes to speak any precedence<br \/>\nthey need for any piece of system reminiscence they need. So long as purposes present affordable hints by way of this extension, implementing<br \/>\nprioritization within the kernel after which using app-provided priorities has the potential to stabilize issues by loads!<\/p>\n<p>Hooking up priorities within the kernel seems to be loads much less of a difficulty than you may count on. The kernel already maintains a Least-Not too long ago-Used<br \/>\nrecord of reminiscence allocations that, on eviction, are traversed so as. For every entry on that LRU record, eviction is tried till there&#8217;s sufficient<br \/>\nfree house for regardless of the eviction was for.<\/p>\n<p>This LRU record offers  heuristic for which utility\u2019s reminiscence needs to be evicted first. Functions that haven\u2019t submitted something in an extended<br \/>\nwhereas are unlikely to want the reminiscence quickly, and since their reminiscence is Not Not too long ago Used, it is going to seem early within the LRU record and be evicted first.<\/p>\n<p>When an utility makes use of a set of buffers, that set of buffers is moved to the very finish of the LRU record in a single bulk. Nonetheless, the order of allocations<br \/>\ninside that bulk shouldn&#8217;t be explicitly managed in any respect. Which means as soon as the kernel closes in on some utility to evict its reminiscence, which particular<br \/>\nitems of reminiscence get evicted is kind of undefined. A simplified visualization may look one thing like this:<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/AMDGPULRU.png\" alt=\"Unsorted LRU list visualization\"\/><\/p>\n<p>If the kernel walks the LRU record like this, it will evict the buffer with a precedence worth of two first, although there are a lot lower-priority buffers<br \/>\nelsewhere within the LRU record. If solely the primary buffer of precedence 2 will get evicted, issues could be okay, but when the extremely essential buffer with precedence 4<br \/>\nfinally ends up evicted as effectively, there are probably going to be issues.<\/p>\n<p>Provided that we already know particular priorities for the person allocations, this LRU record is a quite simple place to combine them. It\u2019s so simple as ordering<br \/>\nthe record entries inside a single utility by their precedence:<\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/pixelcluster.dev\/assets\/images\/AMDGPULRU2.png\" alt=\"Sorted LRU list visualization\"\/><\/p>\n<p>Now, when the kernel goes over the LRU record to search out one thing to evict, the very very first thing it is going to discover and attempt to evict are the lowest-priority buffers.<br \/>\nThe very best-priority buffers are final within the record, and thus solely get evicted when evicting all of the lower-priority buffers was not sufficient.<\/p>\n<h3 id=\"memory-priority-adoption-in-apps\">Reminiscence precedence adoption in apps<\/h3>\n<p>Sadly, not all purposes truly set priorities through VK_EXT_pageable_device_local_memory. As for native Vulkan purposes,<br \/>\nI haven\u2019t noticed any idTech sport utilizing the extension instantly, not less than :\/<\/p>\n<p>The D3D facet appears loads higher, as a result of vkd3d-proton already makes use of VK_EXT_pageable_device_local_memory when obtainable, and interprets each the<br \/>\nID3D12Device::MakeResident\/ID3D12Device::Evict API calls in addition to priorities set through ID3D12Device1::SetResidencyPriority to precedence values<br \/>\nset utilizing the Vulkan vkSetDeviceMemoryPriority command. A number of D3D12 video games make the most of not less than certainly one of these APIs, so the hints these video games present will now be<br \/>\nutilized.<\/p>\n<p>I don\u2019t have tremendous stable numbers for a way a lot reminiscence precisely is overcommitted by most D3D12 apps, as they don\u2019t sometimes expose the full quantity of<br \/>\nVRAM they request in an easy-to-access manner like idTech\u2019s efficiency overlay does. Nonetheless, correctly honoring reminiscence priorities usually appears to have<br \/>\n likelihood to enhance the expertise. Efficiency usually seems extra steady over time (since you\u2019re not counting on luck with which buffers the<br \/>\nkernel evicts as a lot). In some spots I had  comparability level at, I believe it elevated efficiency in comparison with the kernel evicting random<br \/>\nissues by as much as 30% in the perfect case &#8211; however once more, take this quantity with a mountain of salt because it relies upon nearly fully on luck close to eviction.<\/p>\n<p>When all is claimed and finished, how effectively does operating out of VRAM maintain up?<\/p>\n<p>I\u2019d say it\u2019s fairly alright! In lots of instances, it&#8217;s possible you&#8217;ll be shocked how a lot efficiency you may retain even when evicting a gigabyte or extra of reminiscence!<br \/>\nThen once more, that\u2019s after all a fairly optimistic case, and the mistaken factor ending up in CPU RAM can in a short time trigger very vital slowdowns. Eviction is<br \/>\ntough to get excellent, and to an extent, efficiency will all the time be dragged down. If a sport is struggling to hit 30fps even with all the things<br \/>\nin VRAM, needing to evict one thing on prime of all that would typically simply unavoidably end in that 30fps goal being missed.<\/p>\n<p>Regardless, what I hope this blogpost can exhibit is that even when you find yourself with some reminiscence evicted to system RAM, the slowdown will be manageable. There\u2019s<br \/>\nmeasures that drivers (notably, the kernel driver) can take to make overcommit work as quick as potential, and even purposes can do their<br \/>\nhalf in coordinating with the driving force stack to mitigate the results of their reminiscence being evicted. With all the things in place, VRAM overcommit isn\u2019t<br \/>\nactually as large of a deal as one might imagine it&#8217;s at first sight.<\/p>\n<p>All of the work I described right here has already been launched in SteamOS for a while now (it\u2019s each in Secure and Preview. So long as your system is up-to-date, it\u2019s good to go!).<\/p>\n<h2 id=\"a-note-on-upstreaming\">A be aware on upstreaming<\/h2>\n<p>In fact, I\u2019m already engaged on upstreaming all this work so it\u2019s obtainable to everybody! Nonetheless, there\u2019s a variety of transferring elements and a variety of deep<br \/>\nrefactors of some fairly core ideas at play right here, so it is going to probably want time to prepare dinner at the beginning is merged upstream.<\/p>\n<p>On the identical time, I don\u2019t wish to put up a blogpost speaking about a number of cool code simply to complete it with \u201ctruly you may\u2019t see for your self,<br \/>\ngo wait till it\u2019s all upstream lol\u201d, both.<\/p>\n<p>As a center floor, I&#8217;ve rebased the kernel work onto a current upstream model of the kernel and revealed a git department<br \/>\nright here.<br \/>\nWhereas it ought to theoretically yield related results, it didn&#8217;t undergo as rigorous testing the SteamOS kernel did. There&#8217;ll probably<br \/>\nbe bugs and instabilities that weren\u2019t there within the SteamOS model. Use at your personal threat, mainly.<br \/>\nI don\u2019t count on to be sustaining this department in any vital capability, as I\u2019d fairly deal with getting the patches into upstream correctly.<\/p>\n<p>So as to go by way of utility precedence hints to the kernel, additionally, you will want a customized Mesa department I pushed<br \/>\nright here.<br \/>\nRelated concerns because the kernel department apply right here, as effectively.<\/p>\n<h2 id=\"questions-of-my-own\">Questions of my very own<\/h2>\n<p>Whereas I&#8217;d declare to have a reasonably good overview of the driving force facet of reminiscence administration at this level, I&#8217;m not very conversant in how<br \/>\npurposes determine on supplying reminiscence administration heuristics internally, in any respect. I&#8217;d suspect optimizing instances the place you\u2019ve already run out of<br \/>\nVRAM isn\u2019t precisely the highest merchandise on developer TODOs (who is aware of, perhaps the reminiscence shortage is altering that? :P), so perhaps there\u2019s some<br \/>\nunexplored room for efficiency enhancements there?<\/p>\n<p>In the event you, pricey reader, occur to find out about VRAM administration for bigger video games\/engines (particularly when operating out), I\u2019d love to talk! I&#8217;ve a hunch that<br \/>\nthere\u2019s nonetheless perf to be gained by making apps and drivers coordinate higher, however I\u2019m additionally plainly enthusiastic about how issues look from an utility<br \/>\ndeveloper\u2019s viewpoint.<\/p>\n<\/p><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/pixelcluster.dev\/VRAM-Overcommit\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>August 17, 2026 Earlier this yr, I blogged about work I did to enhance VRAM administration for video games. Now, after many months of floating round in mailing lists, the kernel patches are lastly merged upstream and queued for Linux 7.3! Hooray! To rejoice, let\u2019s look a bit deeper at one sentence I wrote in [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3901,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[2554,945,1227,953,4214],"class_list":["post-3899","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-limits","tag-management","tag-part","tag-physical","tag-vram"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>VRAM Administration Half 2: Past the Limits of Bodily VRAM - Future News 24<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"VRAM Administration Half 2: Past the Limits of Bodily VRAM - Future News 24\" \/>\n<meta property=\"og:description\" content=\"August 17, 2026 Earlier this yr, I blogged about work I did to enhance VRAM administration for video games. Now, after many months of floating round in mailing lists, the kernel patches are lastly merged upstream and queued for Linux 7.3! Hooray! To rejoice, let\u2019s look a bit deeper at one sentence I wrote in [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-18T07:51:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-18T09:59:13+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"26 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"VRAM Administration Half 2: Past the Limits of Bodily VRAM\",\"datePublished\":\"2026-08-18T07:51:00+00:00\",\"dateModified\":\"2026-08-18T09:59:13+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/\"},\"wordCount\":5314,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/pixelcluster.dev\\\/assets\\\/images\\\/RDNA3Microbench.png\",\"keywords\":[\"limits\",\"management\",\"Part\",\"Physical\",\"VRAM\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/\",\"name\":\"VRAM Administration Half 2: Past the Limits of Bodily VRAM - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/pixelcluster.dev\\\/assets\\\/images\\\/RDNA3Microbench.png\",\"datePublished\":\"2026-08-18T07:51:00+00:00\",\"dateModified\":\"2026-08-18T09:59:13+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#primaryimage\",\"url\":\"https:\\\/\\\/pixelcluster.dev\\\/assets\\\/images\\\/RDNA3Microbench.png\",\"contentUrl\":\"https:\\\/\\\/pixelcluster.dev\\\/assets\\\/images\\\/RDNA3Microbench.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/18\\\/vram-overcommit\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"VRAM Administration Half 2: Past the Limits of Bodily VRAM\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"VRAM Administration Half 2: Past the Limits of Bodily VRAM - Future News 24","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/","og_locale":"en_US","og_type":"article","og_title":"VRAM Administration Half 2: Past the Limits of Bodily VRAM - Future News 24","og_description":"August 17, 2026 Earlier this yr, I blogged about work I did to enhance VRAM administration for video games. Now, after many months of floating round in mailing lists, the kernel patches are lastly merged upstream and queued for Linux 7.3! Hooray! To rejoice, let\u2019s look a bit deeper at one sentence I wrote in [&hellip;]","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/","og_site_name":"Future News 24","article_published_time":"2026-08-18T07:51:00+00:00","article_modified_time":"2026-08-18T09:59:13+00:00","og_image":[{"url":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"26 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"VRAM Administration Half 2: Past the Limits of Bodily VRAM","datePublished":"2026-08-18T07:51:00+00:00","dateModified":"2026-08-18T09:59:13+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/"},"wordCount":5314,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#primaryimage"},"thumbnailUrl":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png","keywords":["limits","management","Part","Physical","VRAM"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/","name":"VRAM Administration Half 2: Past the Limits of Bodily VRAM - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#primaryimage"},"thumbnailUrl":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png","datePublished":"2026-08-18T07:51:00+00:00","dateModified":"2026-08-18T09:59:13+00:00","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#primaryimage","url":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png","contentUrl":"https:\/\/pixelcluster.dev\/assets\/images\/RDNA3Microbench.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/18\/vram-overcommit\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"VRAM Administration Half 2: Past the Limits of Bodily VRAM"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3899","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3899"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3899\/revisions"}],"predecessor-version":[{"id":3900,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3899\/revisions\/3900"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3901"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3899"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3899"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3899"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}