State of AI with Nathan Benaich - Accelerating science and medicine with collaborative agents
Summary
本期演讲由Google DeepMind研究负责人Vivek Natarajan在Rise 2026大会上发表,主题是用协作式AI智能体加速科学与医学发现。他的核心论点是:把当年造就AlphaGo的自我博弈与搜索机制迁移到科学研究中,关键在于教会模型从快速的“系统一”思维转向缓慢、严谨、深思熟虑的“系统二”思维。AI协同科学家通过“生成、辩论、演化”的循环产生并打磨科学假设,并借助排名智能体进行两两辩论、打Elo分,只把最强的想法交给人类,同时表达自身的不确定性(认知谦逊)。真实案例极具说服力:一位研究者把课题交给系统,两天后它给出了他尚未发表的结论作为首选假设,还多给了四个新方向;在抗生素耐药性、急性髓系白血病和肝纤维化等湿实验室研究中,系统甚至发现了FDA已批准的抗癌药Vorinostat具有抗纤维化活性。Natarajan把这种“机器广度扫描、科学家深度判断”的模式称为互补智能,并坦言目前成果集中在有海量未读文献的生物学,化学、数学和物理更难。演讲的另一半聚焦医疗可及性,介绍了诊断对话系统AMIE,它通过对合成病人的数百万次模拟问诊积累了远超人类医生的“经验”,在《自然》发表的评估中于诊断、共情和建立关系上匹敌甚至超越医生。他强调这不是取代而是增强,最终愿景是把传统的“医患二元”关系变成“医生、病人、AI”的三元协作。
Highlights
-
Two days later, it came back with his unpublished conclusion as its top hypothesis, plus four more, one of which his lab had never considered and is now working on. His first move was to email Google asking whether they had somehow got access to his computer.
两天后,系统把他尚未发表的结论作为首选假设返回,另外还给出四个假设,其中一个是他实验室从未考虑过、现在正在研究的方向。他的第一反应是给Google发邮件,质问他们是不是以某种方式访问了他的电脑。
A startling, almost eerie result that hooks the whole talk -
Real discovery is the opposite. Slow, deliberate, rigorous, the product of chewing on a problem for weeks until the spark comes. He wanted system two style thinking. And to get it, he reached back into DeepMind's own history.
真正的发现恰恰相反:缓慢、深思熟虑、严谨,是对一个问题反复咀嚼数周直到灵感迸发的产物。他想要的是“系统二”式的思维方式,而为了实现它,他回溯了DeepMind自身的历史。
Crisp articulation of the System 1 vs System 2 core thesis -
Most of the team thought it was too early. They did it anyway. Sometimes, as Natarajan put it, you jump off the cliff and then you figure out how to build an airplane on the way down.
团队大多数人都觉得为时过早,但他们还是干了。正如Natarajan所说,有时候你先跳下悬崖,然后在坠落途中才想办法造出一架飞机。
Memorable, vivid metaphor about risk-taking in research -
Vorinostat not only showed anti-fibrotic activity, but cut TGF-beta-induced chromatin damage by over 91%, a hint of regeneration. The interesting part is that Vorinostat is an FDA-approved cancer drug, exactly the kind of cross-field connection a liver specialist might never make ...
Vorinostat不仅表现出抗纤维化活性,还将TGF-β诱导的染色质损伤减少了超过91%,透露出再生的迹象。有趣的是,Vorinostat是一种FDA已批准的抗癌药——正是肝病专家可能永远不会想到的那种跨领域联系。
Concrete, quantified wet-lab payoff showing cross-field discovery -
A human doctor might see 10,000 to 50,000 patients in a career. AMIE has already run hundreds of millions of conversations, building up the world's most experienced doctor. In evaluations published in Nature, Amy matched or beat physicians on diagnosis and more pointedly on rappo ...
一名人类医生一生中大约接诊一万到五万名病人,而AMIE已经进行了数亿次对话,打造出世界上经验最丰富的医生。在《自然》发表的评估中,AMIE在诊断上匹敌甚至超越医生,更引人注目的是在亲和力、共情和建立关系方面也是如此。
Striking scale contrast plus the surprising claim of beating doctors on empathy
Full transcript
accelerating science and medicine with collaborative agents, with Vivek Natarajan, research lead at Google DeepMind, at Rise 2026. Jose Pinedes had spent the better part of a decade working out how one family of bacteria smuggles genes across species, the kind of horizontal gene transfer that helps antibiotic resistance spread. He had the answer sitting on unpublished data, and he handed the same research goal to an AI system to see what it would do.
Two days later, it came back with his unpublished conclusion as its top hypothesis, plus four more, one of which his lab had never considered and is now working on. His first move was to email Google asking whether they had somehow got access to his computer. That story, which Vivek Natarajan told from the rise stage, is the kind of result his team at Google DeepMind has been chasing. Natarajan is a research lead there, working at the intersection of AI, science, and medicine.
When he last spoke at RISE a couple of years ago, the state of the art was MedPalm, a language model tuned to answer medical exam questions. His pitch this year was more ambitious, that the recipe behind AlphaGo can be turned on science in the clinic, and that the trick is teaching models to stop thinking fast and start thinking slowly. System 1 is not enough. The problem with using a chatbot as a scientist, Natarajan argued, is that even reasoning models mostly do System 1 style thinking.
Quick responses drawn from surface level pattern matching. Real discovery is the opposite. Slow, deliberate, rigorous, the product of chewing on a problem for weeks until the spark comes. He wanted system two style thinking. And to get it, he reached back into DeepMind's own history. AlphaGo's 2016 breakthrough came from self-play and search, with agents playing each other, taking feedback from the environment and reinforcing what won.
AlphaZero then showed the same recipe could scale from zero knowledge to superhuman play in months, limited mainly by compute. The AI co-scientist generalizes that idea. Instead of agents playing a game, they generate scientific hypotheses, then critique, debate, and refine them over hours and days. What Natarajan calls a generate, debate, and evolve ideas loop. Borrowing from AlphaStar, DeepMind's StarCraft system, the team added tournaments.
A ranking agent stages pairwise debates between hypotheses, scores them against a rubric derived from the scientist's stated goal, and assigns elo ratings, so only the strongest ideas reach the human. Because the debates run in natural language, they can be summarized and fed back into the agent's context, which is what makes the system self-improving. It also lets the system signal its own uncertainty, what he called epistemic humility, which matters when the scarce resource you are spending is a scientist's time.
The whole project nearly didn't happen. The idea came from Gary Pelts, a Stanford geneticist who, after one of Natarajan's lectures, suggested that a model trained on scientific text might generate hypotheses for the causes of rare disease. Most of the team thought it was too early. They did it anyway. Sometimes, as Natarajan put it, you jump off the cliff and then you figure out how to build an airplane on the way down. From hypothesis to organoid. The penides result.
Run with collaborators at Imperial College on antimicrobial resistance was the moment the team realized they were onto something, but the more telling cases are the ones that ended in a wet lab. Physician scientists at Houston Methodist used the co-scientists to find drug repurposing candidates and combination therapies for acute myeloid leukemia. Pelts' own lab pointed it at liver fibrosis, a disease with few treatments, and tested its picks in human liver organoids. One candidate, Vorinostat, not only showed anti-fibrotic activity, but cut TGF-beta-induced chromatin damage by over 91%, a hint of regeneration. The interesting part is that Vorinostat is an FDA-approved cancer drug, exactly the kind of cross-field connection a liver specialist might never make, and the system surfaced it because it could read broadly while the human judged what mattered. Natarajan called this complementary intelligence, and it is the honest version of the pitch.
The machine goes wide, the scientist goes deep. He kept the limits in view. The system itself is general purpose, he stressed, with nothing in the scaffolding specific to biology. The specialization comes from the tools it reaches for at runtime. But asked where it fails, he was candid that the winds so far have been in biology, where a massive unread literature hides real signal. Chemistry is harder and fields like mathematics and physics, which reward narrow depth first reasoning over wide reading.
Harder still. The common thread is not biology or medicine specifically, but a way of turning compute into discipline deliberation, then putting the result back in front of expert humans. Manufacturing medical experience. The other half of the talk was about access. World-class medicine, Natarajan said, is pretty much a geographic and socioeconomic lottery. And his team's second mission is to close that gap. The vehicle is AMIE.
a diagnostic dialogue system he co-leads with Alan Karthikasalingam, a vascular surgeon still practicing in the NHS. Asked once how to choose between two doctors, Karthikasalingam told him to always go with the one who has more gray hair. There's no substitute for experience, so the team manufactured it. Using the same self-play machinery, Amy ran consultations against synthetic patients and a critic, refining itself over millions of simulated dialogues.
A human doctor might see 10,000 to 50,000 patients in a career. AMIE has already run hundreds of millions of conversations, building up the world's most experienced doctor. With the heavy caveat that it all happens in simulation, it is starting to pay off. In evaluations published in Nature, Amy matched or beat physicians in simulated consultations with patient actors on diagnosis and more pointedly on rapport, empathy and relationship building.
I'm kind of sorry about the doctors and humans in your life, Natarajan, deadpan to the unsurprised. But this is not about replacement," he insisted. The story is still about augmentation. A companion study found that general physicians given complex diagnostic puzzles did significantly better with the AI as a thinking partner than working alone or with standard tools like web search. And in an early supervised feasibility study with Beth Israel Deaconess in Boston.
where patients spoke to the AI before an urgent care visit under physician oversight, zero safety stops were required under the study's predefined criteria, patient trust in AI rose after the interaction, and the system's pre-visit diagnoses held up against the attending physicians, all without the benefit of lab tests. A third person in the room. For most of modern medicine, Natarajan closed. The core unit of care has been a dyad, the doctor and the patient. His bet is that it is becoming a triad.
the doctor, the patient, and the AI, with the machine as a teammate rather than a tool. That is the idea behind the team's next effort, an AI co-clinician. It is a tidy frame, and the evidence on stage made it land harder than it would have a year ago. The deeper claim running under both halves of the talk is that the self-play recipe, which once mastered a board game, can now give scientists and clinicians a new kind of thinking partner, one that searches widely, argues with itself, and hands humans better starting points. The wet labs and the early clinical studies are starting to agree.