← All shows

State_of_AI_with_Nathan_Benaich_Odyssey_raises_$310M_Series_B_for

Published Jun 23, 2026 · Duration 6:34 · Language en · 8 highlights

Summary

本期节目介绍了世界模型公司Odyssey以14.5亿美元估值完成3.1亿美元B轮融资,由Natural Capital领投,亚马逊、GV、AMD Ventures等参投,主播也是其种子轮的首笔投资人。节目核心论点是世界模型不同于Sora、Veo等视频生成器:后者只能渲染固定片段供观看,而世界模型更像模拟器,能持续接收输入、保持持久状态并被实时改变。Odyssey押注足够规模下的下一状态预测会迫使模型学会物理规律,否则长程推演会陷入混乱。过去一年它从三个方向突破:用Star Child 1实现实时音视频同步的多模态体验,用Agora 1构建无需手工引擎的多智能体共享世界,用强化学习智能体Prowl主动寻找并修复模型漏洞。两位创始人来自自动驾驶领域,擅长把不可能的物理AI问题拆解、模拟并使模型泛化。节目还回顾了图像生成从GAN到逼真扩散模型的十年历程,认为世界模型正走上同样轨迹。最终结论是:谁能最快生成最丰富的体验,谁就能在机器人、智能体与仿真领域领先。

Highlights

  1. The same trajectory from pixelated if you're lucky to, I can't tell the difference between reality and AI, is happening in the domain of world models and Odyssey is leading the frontier.

    从'运气好才勉强像样的像素图'到'我分不清现实与AI'的同一条轨迹,正在世界模型领域上演,而Odyssey正引领前沿。

    Frames world models as following the same dramatic photorealism arc as images.
  2. A world model is not a video generator. Sora, Veo, Kling and their successors take a prompt and render a fixed clip. The output can be beautiful, but the future is locked from the outset of generation.

    世界模型不是视频生成器。Sora、Veo、Kling及其后继者接收提示词后渲染出固定片段,画面或许很美,但未来从生成之初就已锁定。

    Sharp distinction between video generators and interactive world models.
  3. By contrast, a world model is closer to a simulator than a renderer. It predicts the next state of an environment from the past and from whatever a participant does next. It holds a persistent state, a memory of the world that can be acted on and changed.

    相比之下,世界模型更接近模拟器而非渲染器。它根据过去和参与者接下来的动作预测环境的下一状态,并持有可被作用和改变的持久世界记忆。

    Core definition: persistent, interactive state vs static clip.
  4. Odyssey is betting that next state prediction, at enough scale, forces a model to learn physics because a model that has not learned physical regularities drifts into nonsense over a long rollout.

    Odyssey押注:足够规模下的下一状态预测会迫使模型学会物理,因为没学到物理规律的模型在长程推演中会漂移成胡言乱语。

    The central scaling bet—physics emerges from next-state prediction.
  5. Nothing is in the intellect, as the line goes, that was not first in the senses. As such, Odyssey's Star Child 1 generates synchronized audio and video in real time. This is the first real-time multimodal world model.

    正如那句话所说,凡在理智中者,无不先在感官中。因此Odyssey的Star Child 1能实时生成同步的音视频,这是首个实时多模态世界模型。

    Bold claim of a first, tied to a philosophical maxim about senses.
  6. Jeff Hawk, Odyssey's CTO, ran a live session of an Agora-1-generated GoldenEye deathmatch with every frame conjured on the fly. Attendees could join the game and play one another in real time.

    Odyssey首席技术官Jeff Hawk现场演示了由Agora 1生成的《黄金眼》死亡竞赛,每一帧都即时凭空生成,到场者可加入并实时对战。

    A playable game with no hand-coded engine—memorable live demo.
  7. Readers may recall OpenAI's faulty reward functions in the wild from 2016, where an RL agent steering a boat learned to rack up points by spinning in circles instead of finishing the race.

    读者或许记得2016年OpenAI那著名的奖励函数缺陷:一个驾船的强化学习智能体学会了原地转圈刷分,而不去完成比赛。

    Famous reward-hacking anecdote reframed as a feature for Prowl.
  8. Odyssey is betting that the next leap in machine intelligence comes from systems that build worlds, act inside them, and learn how reality behaves. The lead in robotics, agents, and simulation goes to whoever can generate the richest experience fastest.

    Odyssey押注:机器智能的下一次飞跃来自能构建世界、在其中行动并学习现实运作规律的系统。机器人、智能体和仿真的领先地位将归于谁能最快生成最丰富体验的人。

    The thesis statement on where the next AI leap and competitive lead come from.
Full transcript

Odyssey raises $310 million Series B for world models. Learning the world from pixels. A decade ago, the AI community was alight with enthusiasm over the first generative models for images, generative adversarial networks. At the time, the model's outputs looked more like a Microsoft paint attempt than anything close to photorealism. Ten years later, latent diffusion models and the scaled-up improved architectures since built by the team at Black Forest Labs have ushered in a level of photorealism arguably indistinguishable from reality where it not for the unreal scenes. The same trajectory from pixelated if you're lucky to, I can't tell the difference between reality and AI, is happening in the domain of world models and Odyssey is leading the frontier. Last week, Odyssey announced that it has raised a $310 million series B at a $1.45 billion valuation led by natural capital. With participation from Amazon, GV,

AMD Ventures, EQT, IQT, and others. We wrote the first check into Odyssey's seed in late 2023. On a personal note, I have known both founders far longer than Odyssey has existed. Oliver Cameron from his years building voyage, and then Cruz, the first self-driving car I experienced thanks to him, and Jeff Hawk from the founding team at Wave, where I was involved from day one. I invited Oliver to speak at RISE 2023 in London.

which created an opportunity for the two of them to spend significant time together in person. They started Odyssey later that year. This round is a good moment to explain what the team has actually been building. The answer is more interesting than AI video. From videos to worlds. A world model is not a video generator. Sora, Veo, Kling and their successors take a prompt and render a fixed clip. The output can be beautiful, but the future is locked from the outset of generation.

As a consumer of the video, your only mode of interaction is to watch. By contrast, a world model is closer to a simulator than a renderer. It predicts the next state of an environment from the past and from whatever a participant does next. It accepts input mid-rollout. It holds a persistent state, a memory of the world that can be acted on and changed. Just as next token prediction enabled language modeling, Odyssey is betting that next state prediction, at enough scale, Forces a model to learn physics because a model that has not learned physical regularities drifts into nonsense over a long rollout. The path to frontier world models. Scale is necessary but not sufficient. The bottleneck is experience. And over the past year, Odyssey has attacked it from three directions. One, making each moment richer. Two, making experience shared and persistent for people and agents. And three, teaching models to generate their own. First, richer experience.

Most world models are mute, which is not only bizarre to experience, but means that a rich information stream is discarded. Sound is where collisions, distance, intent, rhythm, and emotion live. Nothing is in the intellect, as the line goes, that was not first in the senses. As such, Odyssey's Star Child 1 generates synchronized audio and video in real time while responding continuously to streaming text, speech, and action. This is the first real-time multimodal world model.

While this sounds logical, the hard part is that audio and video move on different clocks. A small error in one modality can corrupt the other during a long rollout. Star Child 1's approach is to let each run on its own clock while staying synchronized, turning a bi-directional audio-video foundation model into a causal real-time world model. Second, shared experience. Agora 1 is a multi-agent world model that decouples simulation from rendering.

One function evolves a shared-world state from player actions, while another renders consistent views of that state from independent viewpoints. The result behaves like a game engine with no hand-coated engine underneath. A shared state, the model maintains for every participant, which can be edited into new levels while the dynamics hold. At Rice in London on June 12th, Jeff Hawk, Odyssey's CTO, ran a live session of an Agro-1-generated GoldenEye deathmatch.

with every frame conjured on the fly. Attendees could join the game and play one another in real time. Beyond games, robots need shared environments before they touch the real world. Agents need places to collide, coordinate, compete, and fail. Simulators need to cover worlds that have never existed. Odyssey is pursuing all of these directions. With design partners, it will name in time. Third, self-generated experience.

Prowl is a reinforcement learning agent rewarded for breaking the world model, freezing a waterfall, losing a crosshair, popping geometry under the camera, ignoring a control input, or collapsing through a hard scene transition. Through active exploration, an agent can therefore improve a world model. Readers may recall OpenAI's faulty reward functions in the wild from 2016, where an RL agent steering a boat learned to rack up points by spinning in circles instead of finishing the race.

RL agents are unreasonably good at finding the cracks in a system. Prowl points that talent at the world model itself. The agent hunts a weakness, the model trains it away, and the agent comes back for a harder one. It is a way to manufacture the experience these models are short of. Why Odyssey? When we first invested, we wrote that great research is never enough on its own. The teams that win pair it with execution, product taste, and a feel for where the customer needs to go next.

Oliver Cameron and Jeff Hawke came out of self-driving, where the job is to break an impossible physical AI problem into tractable pieces, simulate the world, make the model robust and generalizable. Odyssey's shipping cadence since their seed round continues to accelerate. Odyssey 2, max for scale and physics, Star Child 1 for multimodal grounding, Agora 1 for shared state, and Praul for closed loop improvement.

Odyssey is betting that the next leap in machine intelligence comes from systems that build worlds, act inside them, and learn how reality behaves. If that is right, the lead in robotics, agents, and simulation goes to whoever can generate the richest experience fastest. We wrote the first check because we think Oliver, Jeff and the team are the ones who will.

Delete this episode?

This removes the episode page and its saved audio from this library.