AI学看世界,不该再靠抽卡
婴儿看世界是连续的:眼前画面一帧接一帧,没有随机打乱,没有反复重播。但今天几乎所有AI的“自学”都像抽卡——把图片随机洗牌、一锅乱炖地喂进去。这篇论文把训练方式改成真正的连续视频流:按时间顺序一帧帧看,不许回头重看,结果发现主流方法全崩了,只有MAE(一种靠“遮住再猜”来学习的模型)勉强撑住。问题不在相邻批次太像,而在同一批里的画面几乎重复——你连续看10帧街景,每帧都差不多,模型就学不到新东西。研究者给出的解法很朴素:别让模型死盯重复画面,主动挑那些有运动的区域去看。改完之后,连续流训练的效果追平了传统洗牌训练,而且数据越多越强。它不是你明天能用上的东西,但它戳破了一个行业默认的假设:AI学视觉,也许该像婴儿一样,老老实实看世界。
📄 原文摘要(英文)
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.