AI Pulse
📄 论文解读

AI 自己当科学家:边猜边验证,发现效率翻倍

现在的 AI 做科研,要么靠大模型瞎猜,要么靠传统统计硬搜,都不靠谱。这篇提出一个「边猜边验证」的循环:AI 每提出一个候选(分子、蛋白质、程序),就立刻用一个小模型预测它好不好,并标出「我有多不确定」,然后带着这个不确定性去改进下一个候选,每做一次实验就把结果记进记忆,越搜越准。在神经网络训练、抗体设计、分子优化三个任务上,它比纯大模型反思和传统搜索都好:抗体结合能降低 18%,分子多目标性能提升 60% 以上。它不是你明天能用上的工具,但这是 AI 从「聊天」走向「自己做实验、自己修正」的关键一步。

📄 原文摘要(英文)

Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新