AI模型存档瘦身33%:把压缩当写程序
AI模型越来越大,存起来越来越贵。这篇把压缩当成“写程序”:先设计一套能表达张量结构的小语言,再让AI自动生成一段能精确还原数据的程序。在10个公开模型上,2.13TB压到1.41TB,比通用压缩器小30%,速度还快。这不是你明天能用的工具,但模型存储成本是每个AI公司都在头疼的事,这个思路可能改变未来模型分发的成本结构。
📄 原文摘要(英文)
Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-specific compressors rely on fixed and format-specific pipelines. We present Brevis, which formulates lossless tensor compression as program synthesis. We design a typed domain-specific language (DSL) that captures recurring tensor structures, such as repeated regions and floating-point fields, through a set of reversible operators. Given a tensor, Brevis synthesizes a self-contained DSL program that reconstructs it bit-exactly. A checkpoint-specific production prior, learned from a small representative sample of tensors, guides a bounded A* search to synthesize compact programs, which can later be executed directly for bit-exact decompression. On 10 public checkpoints spanning language, audio, and image generation models, Brevis reduces 2.13 TB of checkpoint data to 1.41 TB, a 33.93% storage reduction. It produces archives up to 30.87% smaller than those of four general-purpose compressors, including zstd and gzip, and smaller archives than the tensor-specific compressors ZipNN and DFloat11. Under a practical concurrency configuration, Brevis achieves 3.60 GB/s compression and 6.61 GB/s decompression while preserving every source byte.