AI Pulse
📄 论文解读

科学想法也有基因组:AI能追踪思想的进化吗?

科学想法很少凭空产生,它们继承、修补、重组前人的工作,就像生物基因组。但现有AI评测几乎不关注这种继承结构。研究者构建了IdeaGene-Bench,把每篇论文拆解成最小、有证据支撑的“想法基因”单元,并记录它们如何遗传、变异、丢失、引入外来元素。评测包含近2000条谱系轨迹和1000多个基因单元,覆盖10个科学领域。结果:最强AI在谱系推理上准确率仅27.3%,且结构化谱系信息反而打乱了模型排名,并非对谁都帮忙。它不是你明天能用上的工具,但它揭示了一个关键前沿:AI要真正参与科学创新,必须学会理解思想的进化史。

📄 原文摘要(英文)

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新