科学发现源于科学家针对复杂问题提出新颖假设,并经过严格的实验验证。为增强这一过程,我们推出了 Co-Scientist,一个基于 Gemini 构建的多智能体 AI 系统,用于结构化科学思维和假设生成。Co-Scientist 旨在帮助科学家发现新的原创知识。根据研究目标和先前的科学证据,它能够制定出可验证的新颖研究假设。系统设计包括多个智能体持续生成、批评和完善假设,并通过扩展测试时计算来加速。主要贡献包括:(1) 一个多智能体架构,配备异步任务执行框架,实现灵活的计算扩展;(2) 一个锦标赛进化过程,用于自我改进的假设生成。自动评估显示,测试时计算的持续扩展能随时间提高假设质量。尽管具有通用性,我们重点在三个生物医学应用中进行验证:药物重定位、新靶点发现和解释抗菌药物耐药性机制。具体而言,Co-Scientist 帮助识别了急性髓系白血病的新药物重定位候选和协同联合疗法,并通过体外实验进行了验证。这些实际验证展示了 Co-Scientist 加速科学发现的潜力,开启了 AI 赋能科学家的新时代。
Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent AI system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and prior scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system's design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include: (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling; (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. While general purpose, we focus the validation in three biomedical applications: drug repurposing, novel target discovery, and explaining mechanisms of anti-microbial resistance. Specifically, Co-Scientist helped identify new drug repurposing candidates and synergistic combination therapies for acute myeloid leukemia, which were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI empowered scientists.
核心贡献 · Key contributions
提出 Co-Scientist,一个基于 Gemini 2.0 的多智能体 AI 系统,用于结构化科学假设生成。 Introduces Co-Scientist, a multi-agent AI system built on Gemini 2.0 for structured scientific hypothesis generation.
展示了测试时算力在科学推理中的显著扩展,随时间推移提升假设质量。 Demonstrates significant scaling of test-time compute for scientific reasoning, improving hypothesis quality over time.
在三个生物医学应用中验证系统:药物重定位、新靶点发现和抗菌素耐药性机制。 Validates the system in three biomedical applications: drug repurposing, novel target discovery, and antimicrobial resistance mechanisms.
通过选择最高 Elo 评分结果,在 GPQA 钻石集上达到 78.4% 的 top-1 准确率。 Achieves top-1 accuracy of 78.4% on GPQA diamond set by selecting highest Elo-rated results.
在 Elo 评分上优于其他前沿模型(如 OpenAI o1、DeepSeek R1),通过增加算力进行迭代改进。 Outperforms other frontier models (e.g., OpenAI o1, DeepSeek R1) in Elo rating with increased compute for iterative improvement.
随时间推移增强专家“最佳猜测”解决方案,表明 AI 可增强专家科学工作。 Enhances expert 'best guess' solutions over time, suggesting AI can augment expert scientific work.
局限 · Limitations
验证集中于生物医学;对其他科学学科的泛化性尚待证明。 Validation focused on biomedicine; generalizability to other scientific disciplines remains to be demonstrated.
自动评估依赖 Elo 评分,可能无法完全捕捉假设的新颖性或实际可行性。 Automated evaluations rely on Elo ratings, which may not fully capture hypothesis novelty or practical feasibility.
系统需要大量测试时算力,可能限制资源受限研究人员的可及性。 System requires significant test-time compute, which may limit accessibility for resource-constrained researchers.
验证中使用了专家在环指导;未评估无专家输入的完全自主性能。 Expert-in-the-loop guidance was used in validations; fully autonomous performance without expert input is not assessed.
湿实验验证限于特定案例研究;需要在更多样化的问题上进行更广泛的验证。 Wet-lab validations are limited to specific case studies; broader validation across more diverse problems is needed.
论文章节 · Sections(共 24)
摘要Abstract
1 引言1 Introduction
2.1 推理模型与测试时计算扩展2.1 Reasoning models and test-time compute scaling
2.2 人工智能驱动的科学发现2.2 AI-driven scientific discovery
2.3 人工智能在生物医学中的应用2.3 AI for biomedicine
3 引入人工智能共同科学家3 Introducing the AI co-scientist
3.1 人工智能共同科学家系统概述3.1 The AI co-scientist system overview
3.2 从研究目标到研究计划配置3.2 From research goal to research plan configuration
3.3 支撑人工智能共同科学家的专业智能体3.3 The specialized agents underpinning the AI co-scientist
3.4 专家在环与共同科学家的交互3.4 Expert-in-the-loop interactions with the co-scientist
3.5 人工智能共同科学家中的工具使用3.5 Tool use in AI co-scientist
4 评估与结果4 Evaluation and Results
4.1 Elo 评分与高质量人工智能共同科学家结果一致4.1 The Elo rating is concordant with high quality AI co-scientist results
4.2 测试时计算扩展提升人工智能共同科学家的科学推理能力4.2 Scaling test-time compute improves scientific reasoning of the AI co-scientist
4.3 专家认为人工智能共同科学家结果具有潜在新颖性和影响力4.3 Experts consider the AI co-scientist results to be potentially novel and impactful
4.4 使用对抗性研究目标对人工智能共同科学家进行安全评估4.4 Safety evaluation of the AI co-scientist using adversarial research goals
4.5 人工智能共同科学家在药物重定位中的应用4.5 Drug repurposing with the AI co-scientist
4.6 人工智能共同科学家发现肝纤维化的新治疗靶点4.6 The AI co-scientist uncovers novel therapeutic targets for liver fibrosis
4.7 人工智能共同科学家重现抗微生物耐药性的突破4.7 The AI co-scientist recapitulates a breakthrough in antimicrobial resistance