← All shows

State of AI with Nathan Benaich - AI needs science_s search history

Duration 9:50 · Language en · Published Aug 13, 2026 · 5 highlights

Summary

本期节目主张,若要让AI形成真正的科研“品味”,仅学习论文中从假设到结论的整洁路径远远不够,还必须看到科学家曾考虑、放弃和修正的分支。传统论文会抹去不合常规的理论、不可负担的方法、阴性结果与走过的弯路,而这些原始试错记录恰恰可能是训练科研直觉最有价值的数据。剑桥研究者阿利斯泰尔·拉塞尔的团队正在把发现过程记录为动态决策图,但节目强调,日志只有在明确标注失败原因、当时证据、候选方案、预期、置信度、成本、选择理由及后续结果时才有训练价值。由于只有被选中的实验会产生结果,决策史还面临“选择性标签”问题,模型可能学会模仿某个实验室受设备和资源限制的偏好,却无法判断这种偏好是否正确。为避免事后合理化和形式主义,记录必须在结果出现前完成、保留研究者之间的分歧,并以极低的额外负担嵌入日常科研工作。自主实验室和注册报告证明,预先记录意图、快速连接决策与结果能够改善搜索,但开放式生物学还需要资助者出钱偶尔执行“次优分支”,以揭示被系统性忽略的可能性。最终应在不同实验室和领域中比较“只读论文”与“论文加决策史”的模型,以单位成本信息增益、校准度和及时放弃弱分支的能力衡量其科研品味是否真正可迁移;在此之前,科研搜索史仍只是有前景的记录,而不是成熟的训练集。

Highlights

  1. Papers are the happy path. If we want AI with research taste, it has to learn from the branches that were considered, rejected, and never written down. Ideas become nodes in a living graph: they branch when a meeting produces two plausible experiments, merge when separate lines o ...

    论文呈现的是一条顺利路径。如果我们希望AI具备科研品味,它就必须学习那些曾被考虑、被否决却从未写下来的分支。想法会成为动态网络中的节点:当会议产生两个可行实验时分叉,当不同证据链汇合时合并,而当某人悄悄停止探索时则逐渐熄灭。

    Nathan Benaich A vivid blueprint for capturing hidden scientific reasoning
  2. A paper presents a clean progression from hypothesis to result to conclusion. It reads like a browser history with every dead end deleted—the ten open tabs, the four rephrased queries, the wrong turns down a sub-forum—all scrubbed, leaving one clean path from question to answer a ...

    论文呈现的是从假设到结果再到结论的整洁进程。它就像一份删除了所有死胡同的浏览器历史:十个打开的标签页、四次改写的搜索词、误入某个子论坛的弯路,全都被清除,只留下从问题通往答案的一条整洁路径,并被奉为标准路径。

    Nathan Benaich A memorable analogy exposing how papers distort discovery
  3. Claude mythos preview beat the human choice 64% of the time, while Opus 4.5 managed 51%. The study was possible because Anthropic's researchers work inside a tool that logs by default. Biology has no equivalent: its reasoning happens in hallways, at benches, in Slack threads, and ...

    Claude Mythos Preview有64%的时候优于人类选择,而Opus 4.5达到51%。这项研究之所以可行,是因为Anthropic的研究人员在默认记录过程的工具中工作。生物学没有类似环境:推理发生在走廊、实验台、Slack讨论串和白板上,而结果往往数周后才在另一个地方出现。

    Nathan Benaich Concrete evidence paired with biology's stark data gap
  4. Such a record can teach a model to imitate a lab's taste, but it cannot establish that the taste was good. A lab that never ran CryoEM because it did not own a microscope teaches a model that CryoEM is rarely the right call. These must be timestamped before the outcome because hi ...

    这样的记录可以教模型模仿某个实验室的品味,却无法证明这种品味是好的。一个因为没有显微镜而从未做过冷冻电镜实验的实验室,会让模型学到冷冻电镜很少是正确选择。所有记录都必须在结果出现前加上时间戳,因为事后偏见会把不确定性变成看似必然。

    Nathan Benaich A sharp warning about biased imitation and hindsight
  5. Where two branches are plausible and the stakes justify the cost, a lab should sometimes run the runner-up. Otherwise, the record captures what today's scientists usually chose and stays silent on what they systematically overlooked. If decision histories improve choices on unfam ...

    当两个分支都合理且研究价值足以抵偿成本时,实验室有时应该执行排名第二的方案。否则,记录只会捕捉今天的科学家通常选择了什么,却不会说明他们系统性忽略了什么。如果决策史能改善其他地方对陌生问题的选择,阿利斯泰尔的团队就将捕捉到科学界从未成功写下的东西:一份可迁移的科研品味记录。

    Nathan Benaich A bold experiment for proving scientific taste can transfer
Full transcript

Nathan BenaichYou're listening to Science Needs a Search History for AI to Learn Taste on Air Street Press. Papers are the happy path. If we want AI with research taste, it has to learn from the branches that were considered, rejected, and never written down. Actually, that's where the goal is, said Alistair Russell, my graduate school friend who leads a preclinical genome editing group at Cambridge. He was talking about the winding road of science that is omitted from published papers. When you're discussing how to do the experiment, and why this way is better than another way. And what does the data really mean? I know what it shows, but what does it mean? At Ray's 2026, he described how his group has begun logging the twisting path of discovery as it happens, recording the verbal and written exchanges that normally disappear. Ideas become nodes in a living graph. They branch when a meeting produces two plausible experiments, merge when separate lines of evidence converge and go dark.

Nathan Benaichwhen someone quietly stops pursuing them. In his implementation, one agent scores novelty, while another tries to learn how scientists think and how they navigate through a complex world of data so that high potential nodes can trigger deeper investigation. As frontier AI labs deploy agents towards scientific discovery, giving AI authentic scientific taste remains a trillion-dollar dilemma. The bottleneck is that the record we have kept for centuries of science might be insufficient to get us there.

Nathan BenaichThe experiments that never make the paper. Peter Mediwar, 1960. Noble Laureate and father of transplantation asked in 1963 whether the scientific paper was a fraud and answered yes. The form of the paper misrepresents the thinking that produced it. 60 years on, the diagnosis is unchanged. A paper presents a clean progression from hypothesis to result to conclusion. Lost along the way are the unconventional theories, the abandoned or unaffordable methods, and the underwhelming and inconclusive data. It reads like a browser history with every dead end deleted, the 10 open tabs, the four rephrased queries, the wrong turns down a sub-forum, all scrubbed, leaving one clean path from question to answer as the canonical path. What has changed since then is that the discarded material is now worth something.

Nathan BenaichNo high-impact journal wants these artifacts, but raw trial and error is the most likely source of the training data required to develop scientific intuition or what researchers call taste. Earlier attempts to capture the discards were motivated by scientific integrity. The Journal of Negative Results in Biomedicine launched in 2002 to publish rigorous studies that disputed established models or exposed ineffective treatments. Its archive preserved negative conclusions after they had become papers.

Nathan Benaichrather than the live alternatives and arguments that produce them. Biomed Central closed it in 2017, saying the mission had been served now that other journals publish null results. A less generous reading is that in 15 years, it published around 200 papers because almost nobody wants to read a negative result, let alone write one up when it will not count toward academic tenure. AI models are different. Even an uninteresting negative result can be useful, provided it is labeled.

Nathan BenaichEarlier this year, Anthropic put a version of this to the test. The company pulled 129 decision points from real Claude code sessions between January and March 2026 showed models the work up to a human detour and asked what should happen next. Claude mythos preview beat the human choice 64% of the time while Opus 4.5 managed 51%. The comparison was tilted toward the models because Anthropic deliberately chose moments where the human decision had room for improvement. On 127 further scenarios where the human action was already strong, the models improved on it only about 20% of the time. The study was possible because anthropics researchers work inside a tool that logs by default. The reasoning, detours, and outcome are produced in the same working environment. Biology has no equivalent. Its reasoning happens in hallways, at benches, in slack threads, and on whiteboards.

Nathan Benaichwhile the outcome arrives weeks later somewhere else. The record has to emerge from the work itself. Ask scientists to reconstruct it afterwards and we'll create another polished account. A log is not a label. Before any of this becomes training data, a negative result has to say what failed. The first paper published in the Journal of Negative Results in Biomedicine examined 234 negative studies across five leading medical journals. Only 30% discussed statistical power and half clearly defined a primary outcome. Telemodel and experiment failed without telling it whether the assay was underpowered, a reagent had degraded, or the hypothesis was simply wrong, and it will learn noise with confidence. Decision histories have a second missing label too. When a lab considers five experiments and runs one, only the chosen branch returns an outcome. The other four are experimental counterfactuals. Researchers call this the selective labels problem.

Nathan Benaichthe data reveal results only for the action someone chose to take. Such a record can teach a model to imitate a lab's taste, but it cannot establish that the taste was good. It also smuggles local constraints into general lessons. A lab that never ran CryoEM because it did not own a microscope teaches a model that CryoEM is rarely the right call. Alastair's nodes preserve the candidate set, which is the necessary first step. To be useful, each decision record needs six fields.

Nathan Benaichthe evidence available at that moment, and nothing that arrived later. The candidates considered, including the ones dismissed in a sentence. For each candidate, the expected result, the confidence attached to it, and the cost and time and money, the root chosen and the reason. What happened next? How the evidence changed the scientist's view. These must be time stamped before the outcome because hindsight turns uncertainty into inevitability. Disagreement also has to survive.

Nathan BenaichAveraging three scientists into one clean rationale reproduces exactly the information loss we are trying to fix. But there is an obvious failure mode here. Once decisions are logged and scored, people log for the record post facto. Anyone who has watched an electronic lab notebook fill up with retrospective tidying knows how quickly a research tool becomes a compliance exercise. A record that costs a scientist 10 minutes of honest reflection per decision is worth far more than one that costs an hour of performance.

Nathan BenaichA publication initiative called Registered Reports offers a useful starting point. Here researchers submit their rationale, methods and analysis, plan for peer review before the data exists, and the journal commits in principle to publish if they follow the approved plan. Nature has now expanded the format across every field it covers. This approach to paper writing timestamps intent before the outcome, but it freezes one plan. A useful search history must also preserve how the plan changed.

Nathan Benaichwhich alternatives were rejected, when, and on what evidence. Run the runner up. Autonomous labs show what happens when decisions and outcomes are connected in a loop. Liverpool's mobile robotic chemist ran 688 experiments over eight days in a 10-variable formulation space. A batched Bayesian search used each result to choose the next experiments and found photocatalyst mixtures six times more active than the starting formulations.

Nathan BenaichEvery completed experiment changed what the system did next. The robot optimized within a goal and search space that humans had already chosen. It did not decide which scientific question mattered. So what this work demonstrates is narrower, but still useful. A search history pays off when the options return standardized outcomes quickly. Open-ended biology is harder because branches can take weeks and the discarded alternatives may never be run. Some exploration therefore has to be bought.

Nathan BenaichWhere two branches are plausible and the stakes justify the cost, a lab should sometimes run the runner up. Otherwise, the record captures what today's scientists usually chose and stays silent on what they systematically overlooked. A funder could ring fence a small fraction of a grant for the branch not taken, provided that the decision record and the outcome are both deposited. That costs money, but so does having every lab rediscover the same abandoned path. Test whether taste transfers.

Nathan BenaichWhen it comes time to evaluate if our new AI system exhibits taste, the comparisons should run across laboratories and fields. At selected decision points in a live research campaign, we should freeze the six fields, the evidence, the candidates, expected results, confidence, costs, and the choice made. Experts then rank the options before the outcomes are known, and those outcomes are checked independently later. We could give models one of two diets, papers alone or papers plus decision histories, and ask them to rank the next experiment on unfamiliar projects. Then we score information, gain per dollar in week, calibration, and how quickly week branches are abandoned. If history has improved choices only inside the lab that produced them, they amount to useful organizational memory. If they improve choices on unfamiliar problems elsewhere, Alastair's group will have captured something science has never managed to write down, a transferable record of taste.

Nathan BenaichUntil that happens, a scientific search history is a promising record, not yet a training set. Thanks for listening to Science Needs a Search History for AI to Learn Taste on Air Street Press.

Delete this episode?

This removes the episode page and its saved audio from this library.