arXiv:2606.05008v1 Announce Kind: cross
Summary: As multi-modal fashions advance in direction of long-form video understanding, reminiscence emerges as a crucial functionality. Regardless of substantial efforts in creating video datasets and benchmarks, current works primarily deal with notion and reasoning, with out systematically evaluating reminiscence: what fashions retain, how faithfully data is preserved, and the way strong reminiscence stays underneath interference. To handle this hole, we introduce M$^3$Eval, the primary complete analysis framework and benchmark for probing totally different reminiscence dimensions in multi-modal fashions. Grounded in cognitive psychology, our design options fastidiously constructed duties that isolate key elements of reminiscence. Leveraging M$^3$Eval, we conduct intensive experiments throughout consultant multi-modal fashions, revealing constant weaknesses and distinctive behaviors. We discover that fashions battle to take care of disentangled representations when processing parallel video streams, exhibit interference patterns differing considerably from these noticed in human reminiscence, floor reminiscence sources extra reliably within the spatial area than the temporal area, and reveal restricted symbolic reminiscence. Collectively, our benchmark supplies a priceless useful resource for future analysis, whereas our findings spotlight reminiscence as a elementary but underexplored functionality and provide insights for designing more practical reminiscence mechanisms in multi-modal fashions. Our code and dataset can be found at https://pku-value-lab.github.io/m3eval-homepage.
Source link

