AI Pulse
📄 论文解读

AI能听懂120分钟音频,还能精确到秒

现在的AI听音频只能回答“说了什么”,但不知道“什么时候说的”。这篇让AI在长达2小时的录音里,不仅能回答问题,还能给出精确的时间戳——比如“在第45分23秒,主讲人提到了预算”。做法是在音频里每隔一段就插入一个时间标记,让AI学会把声音内容和时间对应起来。在短片段和长录音的测试中,时间定位的准确率都明显提升。它不是你明天就能用的功能,但意味着未来的语音助手、会议纪要工具可以做到“你说哪句,我跳转到哪句”。

📄 原文摘要(英文)

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

arXiv 原文

📬 订阅 AI Pulse

每天三次更新,不错过重要信号

▲ 回到顶部