Choice alignment has grow to be a vital part in enhancing the efficiency of Giant Language Fashions (LLMs), but its influence in Multimodal Giant Language Fashions (MLLMs) stays comparatively underexplored. Much like language fashions, MLLMs for picture understanding duties encounter challenges like hallucination. In MLLMs, hallucination can happen not solely by stating incorrect information but in addition by producing responses which are inconsistent with the picture content material. A main goal of alignment for MLLMs is to encourage these fashions to align responses extra intently with picture data. Lately, a number of works have launched desire datasets for MLLMs and examined completely different alignment strategies, together with Direct Choice Optimization (DPO) and Proximal Coverage Optimization (PPO). Nonetheless, on account of variations in datasets, base mannequin sorts, and alignment strategies, it stays unclear which particular components contribute most importantly to the reported enhancements in these works. On this paper, we independently analyze every side of desire alignment in MLLMs. We begin by categorizing the alignment algorithms into two teams, offline (equivalent to DPO), and on-line (equivalent to online-DPO), and present that combining offline and on-line strategies can enhance the efficiency of the mannequin in sure situations. We overview quite a lot of revealed multimodal desire datasets and focus on how the small print of their building influence mannequin efficiency. Based mostly on these insights, we introduce a novel method of making multimodal desire information referred to as Bias-Pushed Hallucination Sampling (BDHS) that wants neither further annotation nor exterior fashions, and present that it might obtain aggressive efficiency to beforehand revealed alignment work for multimodal fashions throughout a spread of benchmarks.
* Authors contributed equally as first authors.
† Authors contributed equally.

