AI边学边用,自己教自己反而学坏了
AI 在推理时也能继续学习,把新信息写进自己的参数里。但有个坑:如果它学的是自己刚生成的内容,等于拿自己的作业当教材,越学越偏。研究者用 12.8 万 token 的长文本流做实验,三个不同规模的模型都出现同一个现象:保留「自己生成的文本」带来的更新,会让模型在真正的人类文本上预测得更差。问题不在「写作」本身——同样的更新机制用在真实文本上是有益的。他们拆出因果链:用冻结的模型生成训练块,能消除 98% 以上的损害;对比「读到劣质文本」和「在劣质文本上更新」,发现后者额外存下了更多损失。最关键的证据是单次更新:模型对自己刚生成的文本预测更准了,但对新的真实文本预测更差了。这不是明天就能用上的技术,但它揭示了一个 AI 自我训练的根本悖论:AI 一旦开始从自己的输出里学习,就可能陷入自我强化的偏差循环。
📄 原文摘要(英文)
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.