AI Pulse
📄 论文解读

深度不是白加的,残差连接才是关键

把模型做深,一直是提升 AI 能力的默认路径,但新研究给这条路径泼了盆冷水:对大多数主流架构,层数加得越深,效果反而越差,甚至退化。研究者搭建了一个受控基准 DepthBench,在固定模型大小和训练方式的前提下,只改变宽深比例,测试了 10 种架构,发现只有两种残差连接设计(HC 和 Full AttnRes)能在极深形态下持续获益,其余包括常见的 Pre-LN 和各类归一化变体,深度增加带来的收益微乎其微甚至为负。更关键的是,他们通过逐层分析证明,这两种设计的优势来自对额外层数的真正有效利用,而不是其他混杂因素。这提醒我们:堆深度不是免费的,残差连接的设计才是决定深度能否转化为实际算力的核心开关。

📄 原文摘要(英文)

Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (e.g., LayerNorm Scaling) or residual connections (e.g., mHC, AttnRes) to enable better information flow and depth utilization. However, it remains unclear whether they truly translate increased architectural depth into effective computational depth, and whether their reported gains stem from better access to information across depth, or unaccounted-for confounding factors. In this paper, we introduce DepthBench, a controlled benchmark for studying computational depth across various architectures. We systematically vary the width--depth aspect ratio (d_{model}/n_{layer}) from shallow--wide to deep--narrow shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures, we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into improved domain-specific performance and effective computation. Controlled layer-level analyses further show that the gains of HC and Full AttnRes are associated with more effective utilization of additional layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results identify residual connection design as a key determinant of whether depth can serve as a meaningful scaling axis by enabling additional architectural depth to translate into effective computation.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新