AI Pulse
📄 论文解读

开源模型用1/10成本逼近闭源巨头

开源多模态模型Boogu-Image-0.1只用2086万张图和约40万美元训练成本,就在图像生成、编辑、中英文文字渲染上追平甚至超越同类开源模型,并接近GPT-Image-2等闭源系统。它通过改进模型理解、数据质量和训练流程,并在推理时引入智能缩放,实现了低成本下的高性能。所有权重、代码和配方已开源。这不是你明天能直接用的工具,但它证明:闭源巨头的优势可能更多来自系统整合而非模型本身,开源社区正以极低成本快速逼近。

📄 原文摘要(英文)

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新