AI Pulse
📄 论文解读

AI看长视频终于认人了:给每个物体写传记

现在的AI看长视频,靠的是时间轴和文字描述,结果同一个物体换个角度、换个时间,它就认不出来了——你问“那个人后来去了哪”,它只能答出“有个人”,答不出“是那个人”。这篇论文换了个思路:不看时间轴,给每个具体物体建一份“传记”,把不同片段里同一个物理实体的观察串起来,提问时先翻传记、再查证据。在长达一天甚至一周的录像测试里,准确率比之前最好的结果高出4.4个百分点。它不是你明天能用上的功能,但“AI能记住同一个物体跨时间的身份”是视频理解从“看片段”走向“看人生”的关键一步。

📄 原文摘要(英文)

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新