0.8B小模型干翻大模型,文档解析新SOTA
一个只有8亿参数的AI模型,在文档解析任务上击败了所有更大、更复杂的模型。OvisOCR2直接把文档图片转成Markdown格式,连公式、表格、图片区域都能识别。它靠的是数据引擎:用真实文档和合成数据混合训练,再用强化学习调优。在OmniDocBench上拿到96.58分,首次让端到端模型登顶。这不是你明天能用的工具,但说明小模型+好数据也能干翻大模型,趋势信号。
📄 原文摘要(英文)
We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. We build a data engine that combines filtered real-document annotations with synthetic pages whose rendered images and Markdown targets are derived from the same HTML source. The training recipe includes supervised fine-tuning, reinforcement learning on a 4B branch with a multi-component reward design, on-policy distillation into the 0.8B model, and model fusion. On OmniDocBench v1.6, OvisOCR2 achieves a state-of-the-art overall score of 96.58, placing an end-to-end model at the top of this leaderboard previously dominated by pipeline methods and highlighting the potential of end-to-end document parsing. On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06. Beyond these two public benchmarks, we evaluate OvisOCR2 on an in-house benchmark designed to cover a broader set of long-tail and challenging scenarios. OvisOCR2 obtains the best overall performance among the compared methods, providing further evidence of its generalization and robustness. OvisOCR2 is available at https://huggingface.co/ATH-MaaS/OvisOCR2.