arXiv:2607.13639v1 Announce Kind: cross
Summary: We introduce OvisOCR2, a 0.8B doc parsing mannequin. OvisOCR2 is designed as an end-to-end parser: given a doc web page picture, it generates a Markdown illustration in pure studying order, overlaying textual content, formulation, tables, and visible areas. We construct a knowledge engine that mixes filtered real-document annotations with artificial pages whose rendered photos and Markdown targets are derived from the identical HTML supply. The coaching recipe consists of supervised fine-tuning, reinforcement studying on a 4B department with a multi-component reward design, on-policy distillation into the 0.8B mannequin, and mannequin fusion. On OmniDocBench v1.6, OvisOCR2 achieves a state-of-the-art total rating of 96.58, inserting an end-to-end mannequin on the high of this leaderboard beforehand dominated by pipeline strategies and highlighting the potential of end-to-end doc parsing. On PureDocBench, OvisOCR2 additionally achieves the best Avg3 rating of 75.06. Past these two public benchmarks, we consider OvisOCR2 on an in-house benchmark designed to cowl a broader set of long-tail and difficult eventualities. OvisOCR2 obtains the most effective total efficiency among the many in contrast strategies, offering additional proof of its generalization and robustness. OvisOCR2 is obtainable at https://huggingface.co/ATH-MaaS/OvisOCR2.
Source link

