AI模型也有“血缘鉴定”:不看数据,只看权重就知道谁是谁的“后代”
现在开源AI模型可以被随意微调、压缩、合并,但没人记录它们的“族谱”。这篇研究发明了一种“血缘鉴定”方法:不看训练数据,只看模型内部的权重结构,就能判断两个模型是否有共同的“祖先”。原理是,模型在训练时会留下独特的“指纹”,即使经过各种修改,这些指纹依然存在。测试中,它能100%准确地区分有血缘关系的模型和独立训练的模型,而且速度比现有方法快76倍。这就像给AI模型做DNA检测,未来可以帮助追踪模型来源、防止抄袭或恶意篡改。
📄 原文摘要(英文)
Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented. We study data-free white-box lineage verification: can weights alone reveal whether two compatible model checkpoints share ancestry? Residual training produces a shared identity-aligned component in branch products, so this structure alone cannot establish ancestry. We remove it and compare checkpoint-specific structure across residual blocks, yielding a symmetric lineage score calibrated against independent checkpoints. On residual-MLP and GPT-2 benchmarks, the score separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models (AUROC=1.0), distinguishing weight ancestry from behavioral similarity. Under function-preserving checkpoint laundering experiments, weight-space baselines lose margin or fail; our score remains unchanged and runs 76x faster than the nearest robust baseline on GPT-2. The projection-pairing signal appears across six language-model families and beyond, and a case study correctly identifies 3 related and 7 unrelated LLaMA-2 public checkpoints. Collectively, these results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints