← All shows

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) - Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776

Duration 58:40 · Language en · Published Sep 09, 2026 · 6 highlights

Summary

本期节目采访斯坦福教授、BigSpin 联合创始人 Chris Potts,讨论语言学、模型架构、可解释性、AI 使用能力以及“tokenomics(令牌经济学)”之间的联系。Potts 回顾了自己从研究脏话及其语用功能转向 NLP 的经历,并指出流利使用人类语言的非人类系统,为重新审视语言的先天性、统计性与符号性提供了前所未有的实验对象。面对大模型规模化给学术研究带来的压力,他主张研究者不要与资源雄厚的前沿实验室正面竞争,而应探索更奇特、更具创造性的架构、数据机制和解释方法。他认为当前堆叠式 Transformer 极其低效,今天的模型也并非单纯依靠规模获得成功,而是大量关于位置编码、稀疏性、激活函数和量化的工程洞见共同塑造的系统。节目核心提出用类似消费者价格指数的方法衡量令牌的购买力:先定义代码、修复、知识发现等“工程商品”,再比较产出、质量与令牌消耗;初步数据呈现出“tokenflation”,即使模型变好,每个令牌换来的价值仍可能下降。两人也强调不能只评价模型权重,而要评价包含系统提示、推理配置、上下文管理和产品交互在内的完整系统,因为固定模型也会因产品层变化而产生显著不同的成本与结果。关于 AI 素养,研究显示专家倾向于迭代、质疑和修改要求,而新手更常直接委托并无批判地接受答案,因此验证能力和多样化的“对手式”智能体协作比盲目增加令牌更重要。展望未来,Potts 看好递归架构、无分词器的字节级模型以及从数据到能力的因果解释,同时警告少量预训练数据投毒就可能影响模型偏好,学术界应承担高风险、可能改变范式的探索。

Highlights

  1. I did my PhD on, among many other things, swears, what swears are like, why we swear, what information they encode, what kind of taboos exist around them, and so forth. And that was actually the trigger that got me into NLP because I wanted a lot of data of people swearing.

    我的博士研究主题之一是脏话:脏话是什么、我们为什么说脏话、它们编码了什么信息,以及围绕它们存在哪些禁忌。正是这项研究促使我进入自然语言处理领域,因为我需要大量人们说脏话的数据。

    An unexpected origin story for an AI researcher
  2. I feel incredibly privileged to be alive in this moment where humans encounter, for the very first time, non-human creatures that use our language very fluently. From the point of view of understanding the human capacity for language, what a gift.

    我感到无比幸运,因为我们正处在人类第一次遇到能够非常流利地使用人类语言的非人类存在的时代。从理解人类语言能力的角度看,这是一份多么珍贵的礼物。

    Reframes language models as a historic scientific opportunity
  3. The architecture everyone has arrived at, these stacked transformers that we make very deep and very large, are tremendously inefficient. Maybe we could get massively more capable models with half the depth and a quarter of the representational width, and that would be transforma ...

    大家最终采用的架构,也就是被堆得又深又大的 Transformer,效率极其低下。也许我们能用一半的深度和四分之一的表示宽度得到能力强得多的模型,这将彻底改变 AI 的经济性。

    A bold challenge to the dominant scaling paradigm
  4. We've got our tokens spent and we've got our goods produced. Tokens divided by goods produced is a pretty rough measure of the purchasing power of the tokens in those time periods.

    我们知道花掉了多少令牌,也知道产出了多少“商品”。用令牌数除以产出商品数,就能粗略衡量某个时期令牌的购买力。

    Turns vague AI ROI into a measurable economic framework
  5. Experts display an augmentative style. They iterate with the AI. They push back. They complain. They change their requirements. It's really a collaborative mode, whereas novices, low-fluency users, delegate.

    专家表现出一种增强式的使用风格:他们与 AI 反复迭代、提出异议、表达不满并修改要求。这本质上是一种协作模式,而新手或低熟练度用户则更倾向于把任务直接委托出去。

    Explains the behavioral gap behind successful AI use
  6. There is evidence that with very few examples planted in a pre-training dataset, you can have a significant influence on the outlook and preferences and quirks of the final model. So can we detect those examples? What's the nature of those attacks? How well hidden could they be?

    有证据表明,只要在预训练数据集中植入极少量样本,就可能显著影响最终模型的立场、偏好和特性。那么,我们能否检测出这些样本?这类攻击的本质是什么?它们能隐藏得多深?

    Highlights a subtle but consequential model-security risk
Full transcript

We've started to enter a phase of AI that's not just about making models smarter. It's also about making them economically sustainable. As reasoning models consume more tokens, context windows continue to grow, and agents become embedded in more products and workflows, the economics of these systems are becoming impossible to ignore. That's given rise to a new conversation around tokenomics, how we think about the costs, incentives, and trade-offs shaping the next generation of AI.

One person who's been thinking deeply about this is Stanford professor and BigSpin co-founder Chris Potts. His recent work argues that measuring AI progress requires looking beyond model benchmarks to ask a different question. What are our tokens actually buying us?

Here's Chris explaining how he thinks about tokenomics. Another interesting moment to be in as we're all being made aware of the true costs of all this AI usage. The analogy here is like it used to cost me $20 to take a rideshare to the airport Uber or Lyft and now it costs 90. But it's more like 20 to like 500 or something, right? I think what's happening is that the big providers are testing the waters on charging us.

the true costs plus whatever profit they need to make as they all try to gear up for IPOs and so forth. And in turn, that is very quickly leading people to ask questions like what is the return on investment for all these tokens that we have purchased? And it's a very tricky area to be in because what does it mean to think about value in this context? I'm Sam Charrington and this is the TWIML AI Podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI.

through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in. I want to say thanks for coming on. I've been looking forward to this conversation. And I think where I'd love to start us off is to really dig into your background and how it kind of got you, you know, to where you are now. Yeah, my background is in linguistics.

linguistics proper, not even natural language processing. I did my PhD on, among many other things, swears, what swears are like, why we swear, what information they encode, what kind of taboos exist around them, and so forth. And that was actually the trigger that got me into NLP because I wanted a lot of data of people swearing. I wanted to know what the context was like, what their intentions were. So I turned to corpora and From there, you start using NLP tool kits to add structure to those corpora. And then after a few years, maybe you're writing your own tools for doing that work. And then when you look back after 18 years or whatever it's been, you're just an AI person or an NLP person. But that is the true story. And I feel like if I had to, I could trace the lineage of every one of my current projects back to my fascination with why we care when someone drops an F bomb.

So are you a an f-bomb dropper or did you come at it from the perspective of trying to understand these others? I think very infrequently in my life For my linguistics class semantics and pragmatics, which is about linguistic meaning on the final day We always do a class on swearing and I review the history and we kind of tie all the course themes together My handouts for that are full of swears But I only swear once in the lecture. I present the result that people remember things better if the utterance contains a swear because it has a kind of emotional resonance, very primitive reaction. And so in that moment, I pick some fact from the course, some trivial thing, and I restate it with a swear and then I say, all of you will remember this for eternity. But other than that, I'm very shy about it in the class. That's funny. I'm sure they're...

is loads of research on this, but I grew up in New York City, and as a New Yorker, I think that swearing is just kind of part of my natural language and way of communicating, and I married a Midwestern girl, and she doesn't tolerate it at all. She doesn't do it, she doesn't tolerate it, she won't tolerate it for me, and it made for, we've been married for 30 years, so I adapt quickly, apparently.

But for a long time, it took a lot of restraint to change that way of communicating, particularly when I'm communicating about something that I'm excited about or emotional about or want to convey the importance of. It's a really interesting topic.

Well, we're not going to turn the podcast into a podcast about swearing, but I imagine there's enough research there that we could if we wanted to. It's a fascinating area, yeah, because it gets right to the heart of the culture that we've constructed and how it relates to our usage and everything else about us. Yeah, it's fascinating that we have swears. When the old swears lose their power, we invent new ones. We pretend like nobody should use them, but as you say, people use them all the time and it feels like an important part of being a language user that we've...

got them available to us. Yes. Endless string of questions. I'd love to hear your take on kind of a linguist in the age of modern AI, you know, transformers, statistical models. You know, this is a, NLP used to be kind of coming from a linguistic perspective. And now the entire field is shifted to a statistical perspective. And I'd love to hear your reflections on Being on the other side of that transition, as well as maybe more importantly ways that you think that kind of the traditional foundational linguistics is still important to the way we think about AI today. These questions are on my mind all the time because I operate at the intersection of all these different fields. And I will say it's useful to distinguish in this context linguistics and people in my department at Stanford study language and social identity.

historical linguistics, the structure of language, and they're just doing scientific investigation of language as a human phenomenon, and they are not technologists, and they're not trying to inform technology. So their project is interestingly impacted by technological developments. For NLP people who were of course participating directly in the engineering project, they're affected in a very different way by the rise of Gen AI and the kind of homogeneous nature of the solutions that people now adopt in that space.

So for the linguists, I feel like this is the most exciting moment that anyone could have dreamed of. I feel incredibly privileged to be alive in this moment where humans encounter, for the very first time, non-human creatures that use our language very fluently. I think it's weirding us all out, but from the point of view of understanding the human capacity for language, what a gift, because you can ask about the mechanisms.

which are different from humans, but obviously sufficient for achieving a certain kind of behavioral performance. We can think about them as investigative tools. I mean, we train them on the Internet. They're basically incredibly powerful distributional learners, and we can learn a lot from them about the true structure of language by just looking at the kinds of things that they learn. And it really gets at the heart of core questions in linguistics about...

how much of language learning is innate and the nature of our capacity and whether it's statistical or symbolic, all those things come flooding in in a completely fresh way. And so whatever your reaction to language models is, it should be a significant one, right? This should be causing you to rethink key questions. And that's all you could hope for as a scientist that you have new angles, new perspectives, new questions reopened. That's been incredible.

For NLP, I think it's a more uncertain prospect because pre the arrival of like pre-trained models, which for me would be like the ELMO model back in 2017, 2018. Before that, there was still a lot of statistical work, of course, and we were in the deep learning era.

But you could still, for example, do a PhD that was entirely about some specific phenomenon and maybe some very specific tweak to a model. So you could say, I'm going to work on summarization, and I've got a new idea about how to do that well using deep learning models. And that could be your PhD. And what we started to see, 2018, 2019, 2020, especially with the arrival of GPT-3, that that was a very uncertain prospect because you might wake up one morning to find that you had been completely scooped.

that with essentially no effort, one of these large pre-training runs had done better than you at the thing that you'd worked so hard on. And that caused an interesting, probably overall productive, but interesting and challenging crisis for people, especially students who were trying to figure out what to do next with their PhD research. But I think all of us felt a kind of real uncertainty in that moment. Yeah, I remember the anxiety of that time and...

I always felt it was kind of expressed as, you know, is research and NLP fundamentally like scale limited or do you need a certain degree of scale that only a handful of organizations have to do foundational research? And is everyone else going to be relegated to like poking the, poking the beast and seeing what it does? And I'm curious, do you feel Like that was an anxiety that's passed, or is it still very present? Has it panned out quite like that? How do you, you know, how's it been resolved for you? Also fascinating, not resolved. It's something I discuss a lot with my collaborators and with my students. We're all trying to figure this out in this moment. I will say one concrete thing we did was orient a lot of our research toward interpretability.

just the project of understanding how these models end up being so good at such hard tasks. And the reason we did that is it's relatively inexpensive and it's also an area where clearly you would be explicitly hoping that models would get better because then there would be more to explain. Versus if you were doing that summarization project, you might quietly be hoping that there wasn't going to be so much progress so that you could make the progress. Let's hope the next model isn't good at summarization. I want to be the star of that show.

That's, as I said, very uncertain. But if you're doing Mechantrip, you're like, let's get the new model released because now we're going to have even more structure to find, even more to explain. And that felt like a very productive choice. I don't want to leave out the fact that it's also cheaper to do this research and that is significant. And then I would say that right now a lot of us are in a moment of thinking we should do stuff that is weird and creative and out of the mainstream. We should be thinking about trying to achieve the next big thing because competing with these massively resourced, incredibly creative and talented teams is just not a winning game. So let's play a different game and hope that that's, as they say, where the puck is going, not where it is. And what are some examples of that kind of thinking? We've been thinking a lot about architectures because I have a lot of complaints about current architectures. And I would say the other main theme right now for us in my group is thinking about data.

you know, data have strange and wondrous properties. I think we don't understand how data affect models. And that has all sorts of implications for security and safety and also the nature of the learning that these models do. It really data is fundamental. It's all data driven learning. And so telling the full causal story from data to final model state feels like it will just be significant for lots of questions. But I wouldn't want to leave out the architecture one because I feel like the Architecture everyone has arrived at, these staffed transformers that we make very deep and very large are tremendously inefficient. You would hope they were using all that depth and all that representational power to learn modular recursive functions for things and all sorts of exciting stuff. It is not what we find and that seems like a real opportunity to just level up and do better and maybe we could get massively more capable models with half the depth and

a quarter of the representational width and that would be transformative for the economics of AI in addition to leading to all sorts of exciting things for capabilities. It's funny and maybe a bit validating for me to hear you say that because whenever I articulate a thought in that direction, particularly with folks that are coming from the frontier labs or essentially the frontier labs, I get back this kind of feeling that, yeah, you're just not bitter lesson piled enough, like structures, that's old school thinking, you know, you're just trying to like train some features, just collect a lot of data, throw it at the model, and that's all you need. Okay, but here's my response to them. Let's say rewind to 2017, we've got the transformer. It's got absolute positional encodings.

And it's got a particular structure for its MLP layer, which is pretty narrow and pretty dense, and a certain structure to its activations and its layer norms. That's 2017. The bitter lesson-pilled thing to do would be to scale that up. But just consider, for example, how much it would cost to use the n squared attention and the absolute positional encodings would have a context window of 1 million. This is the bitter lesson-pilled thing, right? Just keep scaling.

but it would be absurd. It would cost trillions of dollars to produce models that we all interact with right now. What did people do instead? They thought hard about locality and they thought about how like positional encoding should be favoring local relationships. They completely rethought the MLP so that it's now wide and sparse. Everyone did careful work on the activation functions to make sure there weren't weird outliers so that they could quantize in a good way.

And so forth and so on, all of this analysis work built on intuitions about data and learning led to the model that we have now, which is like a ship of Theseus compared to the 2017 transformer. The only thing that survives is attention and the feedforward layer. And I claim for you that none of that stuff is bitter lesson-pilled. That was all analysis work that was meant to save based on priors and the data and priors about how they knew learning would happen.

So I go back at them. You're not bitter lesson-pilled enough, apparently, although this is a reductio, I think. Oh, I love this. That's such a great response. I think it also really calls out the relationship between data, Mech and Terp, and efficiency, like core themes that you've been focused on and how they interrelate and support one another.

Yeah, absolutely. And this relates to one of my hot takes. You know, it's very fashionable, especially among interpresearchers, but I think in general, for people to say, we don't understand how these models work. It is also very mysterious to us. But the truth is that people in the field have very deep intuitions about how these models work. And that is a causal factor in us making so much progress, because they could think analytically, what would the structure of positional encodings and attention be so that I could do this at million context scale?

You can only achieve that kind of thing based on deep analysis and insight, not by just guessing. And so when people say, oh, we don't now understand, I say, I think you understand much better than you're letting on. I think you understand at least as well as my car mechanic understands how my car works. There are mysteries, but you can take a lot of action and be very effective in improving things. Why do you think they say that? Why do you think they say that they don't understand the models? There's got to be some payback there.

It's probably a paradox of expertise, right? So the more you do know, the more you feel like there are also mysteries, and it's hard to step back from that and be objective and say, yeah, well, we did make a phenomenal amount of progress, and that can't be just because of happenstance. That was because we know a lot. But all you see as an expert is all the things that are still to be explained. Partly also, it's just a narrative in the field, and it does stretch back to days when I think we had very little understanding of how these models worked.

possibly because a lot of them weren't that good, there was very little to explain. And so that's just been slow to catch up with how much progress we have made in understanding the kind of intuitive human level mechanisms that these models are operating with. I also wanted to ask you about DSPY. I forgot about this as we were talking earlier, but you were involved in DSPY, which, well, I'll let you talk about it, but I'm curious how it...

connects into your research and like, you know, some of these pillars that we've talked about. Oh, there's lots of wonderful strands. And what a meta strand I could offer you because we were talking about being strategic with research. This does stem from Omar Khattab, my student. He's the visionary behind DS Pi and still its lead. And he just had the intuition early on that we should rethink what it means to make a scientific contribution.

Previously, we saw it in terms of papers as the beginning and the end of all of this kind of thing that you would contribute. We should instead, he said, think about projects and about empowering people. And so for him, the paper is one part of a broader contribution that might actually be centered on an open source or open weights release that would allow people to do big things. And that's where you find impact. And that's the nature of a contribution going forward.

And DS Pi is a kind of embodiment of that, although he made a similar investment with the Colbert retrieval model. And then people built on what he did. And then you really saw it take off, where open source contributions made it easier and easier to use that technology, leading to more and more impact. And of course, DS Pi is another wonderful example, because in investing in this community, and in the open source resource itself, he built a huge following.

There are lots of startups, mine included, where the core tech stack for the LMS is built on DSPy, and that has made life so much easier. And then, of course, it was a platform for him and for us to really think in an innovative way about prompt optimization and agent work flows and all of those things. Yeah, I was thinking not too long ago the degree to which model strength.

as a Corolla to model size, I suppose, and capability has kind of overcome the need for an explicit framework like DSPY, DSPY. Yeah, there are kind of two levels to that. The one would be just the engineering side where DSPY is great four years ago because it's kind of hard to construct the code around one of these systems in a way that's modular and reproducible and so forth because pecking out something where you've got a prompt string in the middle of your code with some slots in it. It's very error prone and it leads to bad system designs. And DSPy solved that. And you could think that the need for that is diminishing somewhat because now we all specify these systems in English and have the coding agents do them. And even before that, there were, you know, another hundred frameworks that solve that particular part of the puzzle. Oh yeah, there's always competition. And I think at that level of just thinking about

programming interfaces and APIs, they can all learn from each other. And so DSPy learned a lot from PyTorch in terms of layer-wise design and the kind of modularity that introduced. And then, of course, you would hope that everyone kind of slurps up all these interesting innovations and it leads to everyone being better. There's lots of evidence of that at the level of interfaces. I would maintain for you that even if we have agents actually writing the code for these systems, it's great for us and for them if they write it in something that actually expresses these systems as modular components.

so that we can audit them so that they can change them. It just feels like good engineering practices for any agent to think in a modular way. And that's what DSPy encodes. The other side is like the prompt optimization side and a belief people have that the need to be careful with your prompts is diminishing over time. I understand that narrative, but people should also, for example, just run like a simple annotation study.

where they use a few different models or the same model a few times on slightly different data, they will be blown away by the amount of variation that still exists. To be charitable, let's say that these LLMs disagree about fundamental facts about how to label certain texts or what kind of response to give. We all kind of slip past this because we feel like, hey, they're smart and they're good and they're getting better. But if you quantify it, it's pretty disturbing. And the next step from that is to think about...

having all those agents optimize a prompt so that their behavior is at least consistent. And then you're right back at that DSPy vision. It's interesting that you say that because I don't feel like that necessarily aligns with my recent experience. And in particular, one thing that I've noticed that's been surprising is how well aligned I guess.

Maybe that's not the right word, but how similar the responses I get to query across different models. So for example, you know, these are often kind of what I would call like a casual prompt, a casual query, something that I might, you know, type into Google and it will now generate an LLM response for me and it's kind of AI mode.

and I'll take the same thing and put it into chat GPT and maybe Claude. And it surprises me that the responses are often very, very similar, like very similar structure, very similar facts, very similar citations. And stepping back, there are lots of ways that they could answer or approach these different questions, but it seems like...

the models or the training or the system prompts or something is all kind of converged on something that makes the models express themselves very similarly, which seems to be at odds with the last thing you said about the need to optimize prompts or the impact of the individual prompt.

I'm open-minded, but for example, like we just did it, we did a thing recently, we were writing a grant and we needed a title and you want to be strategic with these titles. So we come up with a whole bunch of them ourselves and then we all disagree on what would be the best. So let's find out what the agents think. So ask a few anthropic models and a few GPT models, which of these five titles, which is the best? So you get a different answer from all of them.

along with a detailed rationale about why obviously, of course, the choice that the model has made in that moment is the best one. This is great because then we can think about which one of these arguments is most persuasive. But if you were hoping for consistency at a subjective labeling task, which this is one, you can see right there that you're going to have a real problem unless you give very specific criteria. And then you're kind of also constructing a prompt for them. And you might want to manage them differently. There is a real...

I don't have evidence for this yet, but we have an intuition at Bigspin and the research we've done that you get a kind of paradox that the more requirements you add, actually the more variation you'll see because the different models will key into different subparts of the requirements. And since they do it very concertedly, you can actually get systematically biased behavior from something that you thought was a very good specification. And that again calls for this idea that what you need to do is figure out what the labels ought to look like.

and then have some automatic optimization process get the model there. And that's what things like JEPA and Meet Pro were for. A topic that I really wanted to, a topic that I would really like to dig in to with you based on our previous conversation was the idea of tokenomics. It's something that people are talking about a lot recently.

I think, you know, folks that use cloud code, for example, have like a very visceral experience with anthropic changing the terms around uses, but it's happening under the covers with all of these large providers. And so I think way more now than, you know, six months ago, like we're all...

a little antsy with the relationship we have with these big model providers and the value that we get. You recently wrote an article about this. Talk a little bit about how it ties into kind of your broader research, but also some of the things that you found when you started to dig into this area. Yeah, another interesting moment to be in.

as we're all being made aware of the true costs of all this AI usage. I saw a tweet from Ed Zitron, just a screenshot from someone who was noticing that co-pilot was telling them that their bill last month was $500, and if they keep up the way they are with co-pilot's new billing, it will be $11,000 in the next month, which is real sticker shock.

And the analogy here is like, it used to cost me $20 to take a ride share to the airport, Uber or Lyft. And now it costs 90. I use that analogy as well. But it's more like 20 to like 500 or something, right? Right. If only the slope will be as shallow as we were, right? That's right. We start to wish for those easier stories. Yes. And so what will happen? I mean, I think what's happening is that the big providers are testing the waters on charging us.

the true costs, plus whatever profit they need to make as they all try to gear up for IPOs and so forth. And in turn, that is very quickly leading people to ask questions like, what is the return on investment for all these tokens that we have purchased? And it's a very tricky area to be in because what does it mean to think about value in this context, even if we focus in on people who are doing just coding with coding agents?

Can we agree on what it means to add value and maybe we have a few measures in mind like making a pull request or Committed lines of code that last in the repo for a while or Documentation touched or skill files created But we might also worry that that's not Capturing the value for many kinds of sessions we have which are more open-ended and about discovery So that's the first question is just solving this value issue, right? Let's just agree on what it would mean to add value for a for a coding agent. I like that this line of inquiry because to me, it's the response to this thing that drives me crazy, which is, oh, big tech company CEO. This year, 95% of our code will be generated by AI.

Yeah. Hey, what does that really mean at that level? Like what's what's what are the details beneath there? But is that a good thing? Oh, another dimension, right? Which is is that code a liability or an asset? Right, right. And this idea of like you articulate as kind of code longevity in the code base. That's an interesting way to think about it. There's probably a lot of interesting ways to think about it that very few are thinking about right now.

and all of these fall victim to the standard thing that once you make it a metric, it's no longer useful to you. If we said, oh, it's completion of projects, right? Well, then everyone would just have many projects that they completed, but they could all be liabilities and add very little value. But one framework we could offer that we did in the research you alluded to is, let's think about this like economists might. We might have like a consumer price index. And the first step will be, what's the basket of goods that we're going to consider?

In that standard land it would be like the price of eggs and the cost of rent and other kinds of tangible goods. What are engineering goods that we might track? Eggs might be a summary or a pull request, a bug fix or something like that. Or we could think broadly because we both use these coding agents. Requirement discovery, knowledge accumulation. These are things that we don't currently track, of course, even as engineers.

but might be behind our intuition that these coding agents are making us productive, even if it's not reflected in the PR counts or whatever, right? I mean, in a sophisticated approach, you might say, I don't want more PRs because this is just a certain kind of busy work that doesn't relate to the actual goals I have. What are the actual goals? It's completing valuable projects and so forth. If I could do it with fewer PRs, but I had really robust code, I'd be possibly happy with that.

So we got to figure out what the basket of goods is but then we could start to track it relative to token usage and that would be the consumer price index. So for any time period we could just say I've got my tokens spent and I've got my goods produced. Tokens divided by goods produced is a pretty rough measure of the purchasing power of the tokens in those time periods. Then you would do the standard consumer price index thing of making what they call a hedonic adjustment. So you could just say maybe quality is improving over time so you'd pick some measure for that and make an adjustment to the line. And when we did that study, we did code survival. So the number of lines of code that survives more than four days in the repository, we made an adjustment upward because that rate is going up. That's surprisingly short. Four days. Four day survival. So again, all this is around measurement and I'm happy to just be starting this discourse because we can see it's important to the economics of AI and it seems like the work isn't being done.

at a high enough rate for us to get a clear picture. So we could make it longer and maybe the adjustment would be different. I think currently for the data we have, which is this sweet chat benchmark, which was released by researchers at Stanford, it's about 6,000 real coding sessions, all the metadata, everything you'd want. What we see with Opus 4.6 usage in the time period we have, which is February to mid-April of this year, a decline in the purchasing power of tokens that CPI is going down.

And again, I just want to open the question. Is it because we have the wrong basket of goods or is it because we're actually getting less value from these tokens? The one thing I can say that's kind of definitely a causal factor here is that in February of this year, most of the tokens went to producing code, which relates to the outcomes we just talked about. By mid-April, it was quite split between code generation thinking.

and also explanation to the user. And so that split now is going to have an effect on the things we're measuring, and that might be caused for reflection. There's value in those explanations that's not reflected in PRs, but might be reflected in something like knowledge discovery. And I see that coming up within the same timeframe. It's become very common to now talk about the token efficiency of a new model that's been released with the implication being tokens of internal use tokens, thinking tokens versus per token of output, I guess, is maybe a way to think about it. Another fascinating dimension, and this actually relates all the way back to the theme of efficiency for these architectures. So here's a claim I'll make for you. Based on my read of the literature on inference time scaling, what's sometimes called test time scaling, which is just having the models generate lots of tokens at the moment that you ask them a question. So those scaling trends.

Everything we're seeing now is completely in line with those predictions, which is you get pretty good gains for a while with the more tokens you spend on a log scale. So this is jumping up quite a lot, but you do see it reflected in performance improvements, but it flattens out over time. And it's not like this curve skyrockets. It's sobering. You got to spend a lot of tokens for small gains in performance. We all knew this.

We all do this and we're just seeing it now play out. And when people talk about token efficiency and worry about this, I think what they're seeing is just the real lesson of what we already projected from inference time scaling. And this is independent of the approach to inference time scaling you're taking, whether it's multiple parallel inferences or some kind of Oracle or any number of other schemes. It's just fundamental to inference time scaling.

That's a great question, right? I think we know that it's independent of some of those things, like the parallel work versus having it do lots of long chains. But some of the other factors you mentioned, I think we just don't know. And that's why I said it relates back to the question of efficiency for these architectures. If we made a fundamental change to how the models work, maybe these trade-offs would be very different. I mean, after all, so all of this stuff is a kind of patch job on the fact that there's no recursion in the depth.

It's a fixed depth. And so the only recursion we can get, the open-ended notion of computation, is by generation. But if we had models that could be recursive, maybe fewer tokens for larger gains, I think we don't know. Yeah, I mean, in the end, we're going to spend the cost on compute or tokens. So this might not affect our bills in the end, but it is a fascinating question. What are the true scaling laws and what's possible in this space?

And you're right to push back. We talk about these things like they were like platonic ideals of laws, scaling law invokes that, right? But even for the scaling laws for pre-training, you know, there's lots to discover there. And many of the stories of progress are actually like transcending the scaling law. We see like better improvements than those laws predicted because everyone worked so hard behind the scenes to do very innovative things, which maybe relates to our bitter lesson discussion. Any particular example come to mind of that?

Data usage in the nature of the data really matters. And overtraining the models really matters, which is kind of pushing up against the standard scaling law presentation. And now I'm just going to speculate. I should check on this, but things like mixture of experts might have really flipped the script on what it means to count parameters and then turn how these laws relate. And then I think maybe even also stuff like the context window and so forth. This is another thing to check, but I just speculate that we've seen larger gains from pre-training than you would have predicted by those early scaling loss papers, suggesting that there is some innovative thing that was happening on top of pure scaling. Thinking about the concept of a market basket, one kind of pushback that comes up for me is in the real economy, eggs is different than milk, is different than beef, etc., etc.

And they're all influenced by different factors, you know, production, for example. Whereas what you've done with this kind of CPI basket with tokens is kind of like more like analogies, like here's the typical bundle of work and, you know, what it requires from a consumptive perspective, but.

The tokens aren't fundamentally different. They're the same tokens. It's just how much it takes to do this versus how much it takes to do that versus how much it takes to do that. Tell me what I'm missing there and what does characterizing these products give you in your analysis? Fascinating to think about. One thing I could insert there is the tokens are different at the level of being used for code generation or...

skill file writing or explanation or thinking, right? Those are different kinds of tokens that probably do feel tangibly different to us. So is that an element in your thinking? I think I was thinking from our conversation that you had 10 different, almost like tasks, like 10 different types of tasks from the domain of code generation, which, you know, if they were all kind of largely code generation, you know, that is the part that had some dissonance for me. But if you're talking about Like if your basket is like creative writing versus a few code generation things that are kind of in different versus summarization versus editorial commenting feedback. Those may be more fundamental. Yes. So I think this is very significant. When we have done some research on this as well at the level of what kinds of session types exist and in turn what kinds of users are there.

So you might notice of your own behavior. I guess this is reflected in your comment that sometimes you want to quick check in on a question. Sometimes you want a quick PR to get fired off. Sometimes you want to be in a mode of deep collaboration. Sometimes you're partnering with the AI. Sometimes you're delegating the work and so forth and so on. And the outcome measures that we choose should be sensitive to this. We shouldn't penalize the agent if your chat interaction with it about some scientific question didn't lead to a PR.

was never on the table in the first place. Whereas if you're trying to get some work delegated that's actually a coding task and all it does is chat with you, that would feel quite unproductive. So we need to bring that in and that would be a higher level discovery process of what people are trying to do and so forth. In thinking about the notion of value, is this something that you're anticipating?

Like it strikes me that that's an entire, you know, research thread that, you know, one could go into, I don't know if that's a linguistics or a linguist or an economist or a computer scientist, you know, probably interdisciplinary, like most interesting questions. But is that, you know, is that something that you're working on or was it something that you put out there for someone to take up and run with? I am not sure. I can tell you the...

The lineage of this idea is that we founded this startup BigSpin because we would like to see more people benefit from AI. Whether you love it or hate it, it's here. And I would like the benefits to be more evenly distributed. And I can tell that that will mean bringing on board many more people than currently benefit from AI. Right now, I would say that it's mostly experts deriving real value. And a lot of the world is currently even trying to figure out what this is all about as a tool or an entity in their lives. So we would like to have more access and more productivity. And that implies making the user experiences much better. Figuring out what interactional patterns lead to success for people, meeting them where they are in a kind of adaptive way. The whole list of things that you might worry about if you were a product manager who had some deployed AI product. And I think by that route and from that perspective, we just ended up

worrying about our own token usage increasing and wondering whether there's real value there. And it just happened to collide actually just like three weeks ago with this emerging narrative on the back. I think of all these rumors about IPOs about what the return on investment was. And then all these CEOs came out and said, oh, our spend was enormous and we want to scale back. And we're walking back our claims from a few months ago. And that is just a fascinating thing to witness in the narrative here. I'm wondering, are you also Does this research also attempt to project forward? In theory, you could create a model for anthropics cost and spend based on publicly available data and some presumptions and give us a sense for how close we are to paying full freight for our tokens versus if we're only paying 10% for our tokens.

You could then project what that cost might look like over time as we're paying more and more of the full cost. Again, fascinating questions. I don't have resolving answers. I am glad I am not tasked in some organization with projecting spend on all of this stuff because I think it would be basically impossible. For the time period that I was describing for our little CPI experiment, Anthropic changed the default reasoning on the model at least two times.

So we see they launched it with default reasoning high. We have a mysterious sudden rise in the token usage, which we cannot explain. And there's a new baseline. They lowered it to medium as the default. They patched a bunch of bugs that were related to context management and then turned it back up to high. And all of these things have an effect on the total token output. As you can imagine, they also changed the default context window, which meant people could swallow up much more stuff at any given moment.

So imagine trying to predict what token spend is going to be like when you have all these exogenous events in addition to changes that we don't even know about and questions about where the value actually lies. Very difficult. And then, you know, the true cost of a token, the estimates very wildly for every dollar we spend, it could be as low as two and as high as 20. And I think this is just because it's hard to factor in things like R&D and future build out and depreciation and all of that stuff.

I think at the current moment, we just don't know, but there couldn't be a more significant question for the global economy, basically, than where the value is and who's going to pay and how much. Yeah. In your article, you coined the term tokenflation to describe at least the recent behavior of token economics. I imagine you see that continuing. Seems to be continuing. Yeah.

Yeah, that's certainly the picture that we get from the CPI, a picture of tokenflation. Yes, your token is not buying you what it once did, according to everything we can think to measure here. And even adjusting for models getting better, right? That's critical there. Because if it was just a story of models thinking more and being more robust, and we were all getting exponentially better outcomes from this, then the spend would look completely rational. But that's not the picture that we see. And so we have to do some hard thinking about what's going to happen and how to improve the situation.

Let's dig into that a little bit more. How would you articulate what you're seeing? The models are getting, quote, unquote, better. There's a set of open questions about, are the reported ways that models are better actually reflective of some intrinsic betterness? And that question brings...

It is often about like benchmarking and learning the benchmarks, overfitting that kind of thing. And then there's the kind of question of chattiness and the volume of thought that it requires a given model generation to produce an answer. What are other factors that you see? Yeah, we can pick that apart as well. And this relates to...

a line I've had consistently, which is that we should think in terms of systems, not in terms of models. So in the data that we've got, opus and sonnet 4.5 versus 4.6, those two generation changes, those are real model changes, I assume. I think they did something very substantive at the level of the weights. And everybody immediately saw that that led to like a 5x increase in token usage. And this was related to the introduction of adaptive thinking.

Now, fix that. That's the level shift that we already took, and maybe we're seeing improvements there that are worthwhile. It gets hard to say, but let's assume there was a level up in improvement. Then for the period that we did our CPI experiment for, that's a fixed model, OPUS 4.6. So all the code improvements that we saw in the data relate to the product. This has to relate to things like them turning the knobs on the adaptive thinking, changing things about the system prompt.

changing things at the level of the product. And that's where the improvements were. And so that shows you that even for a fixed model, we can get very different outcomes for these things because they really are sophisticated engineered systems at this point. Yeah. And so was the product in this case specifically cloud code or? Oh yeah. And so we don't, there's tons of stuff there. Yes. I believe we know that these are all cloud code sessions that we kept in our data. Sweet chat is broader than that and involves a couple of other coding agents. But I think I can say, that all our data are cloud code sessions using Opus 4.6. Have you seen any evidence that changes via API usage experience, similarly dramatic variation in performance? Oh, fascinating. To kind of control for a lot of that product level stuff, all the prompts that are hidden from us, all of those affordances. I don't know, but that's a nice thing to think about because it gives us a...

more things that we can control for and more things that are knowable. So kind of in parallel to the model evolution, there's also evolution of the user. You've alluded to this a little bit about kind of your concept is that most AI users now are experts. Talk a little bit about the role of expertise. I think this is also kind of Echoing back to our conversation about DS Pi and like prompt optimization You know you've done some research into how folks are using these models and the role of you know AI fluency tell us about that research. Oh, yeah, first I should say so the Distribution of users across expertise levels I So I guess the nuance picture I'd offer is that the people deriving a lot of value from AI in the current moment tend to be experts

It must be the case that most users of AI are beginners just because the numbers are so large and expertise can't be that widely distributed yet. And that's a very interesting thing because I think probably most things are getting designed for those experts implicitly or explicitly. But for the whole economic picture to work out, many more people need to derive value from this via one avenue or another. And so that does shine a light on this expertise thing as a real factor.

And the headline result there actually builds on something that Anthropic did. They have this AI fluency index. And their core observation in that work is that experts display an augmentative style. They iterate with the AI. They push back. They complain. They change their requirements. It's really a collaborative mode, whereas novices, low-fluency users, delegate. So they trust in the AI. They let it do its thing.

They accept the responses uncritically. And our contributors just show that this is a causal factor in success with these products right now. Experts can do harder things more reliably as a result of all that friction they introduce, all that pushback. Whereas novice users, they accept, but they end up accepting the wrong thing and they're not able to level up from the basic tasks that they think to start with.

That's obviously significant and it feels so tantalizing because pushing back is a natural human behavior. I feel like we could encourage everyone in the world to do this. We probably need to get them out of the mode of thinking it's a super intelligence. You should just trust it. That has been the narrative for a while. What we're seeing in the current moment and possibly for the foreseeable future is that you got to complain, collaborate, introduce yourself, push back, all that stuff that I think we do that we take that for granted, right?

Yeah, yeah. And so from a methodology perspective, how did you approach exploring this? We built on the work that Anthropic did, which they set up a nice framework with some independent research who were doing this kind of usability stuff. And we just have an annotation protocol. We can talk in detail if you want about this, but at BigSpin we have lots of these best practices around having language models essentially collaborate on annotation projects to kind of triangulate on the truth and factor out their individual biases. So we do that stuff and we apply all these fluency markers and then separately we do a thing of estimating task complexity and looking for signs of visible and invisible failures. And so it's the connection between the fluency markers and the task complexity success metrics that was our contribution there. And that's where you can see.

high-fluency users are the ones doing harder tasks. Paradoxically, there's more signs of failure for them because they complain, they push back, they're trying harder things. But as part of all that friction, they're successful with harder things as well. And have you were to try to apply this insight from the perspective of someone in an organization that's trying to help or guide their organization to be more successful?

with AI, like what do you think are the key lessons of this fluency work? If it's an org that's just starting out and wants people to figure out how this could be part of the organization's mission, it would just be that pushback message. And you could do an experiment where you interact with it about something where you're a world expert. We're all an expert in something. Engage in a discourse was one of the best models about something you're an expert in and see how often you feel you have to push back. And this could be a kind of lesson.

for other spheres where I don't know the answer. It might be just as errorful. That could be a good visceral thing. If the org is very far along, I think the main thing to do right now is to have a team of these LMS interacting to improve things. For example, at BigSpin, I didn't set this up. Our founding engineer is very future forward on agents, and he's incredible at this. And when we do PRs now, the first round of review is the agents all interacting, collaborating, disagreeing.

they do the first round of comments, they do the first round of code updates, only after they've resolved things do we look at a PR. So the final human stage should be very high value and the agents did all that work. But when you have one agent do it, they often just reinforce themselves and you don't get good outcomes. It's that team of rivals thing that is transformative. You know, I think it's interesting because, you know, on the one hand, like, of course, that makes sense. But on the other hand, there's something And it also implies that you shouldn't be using these things in areas where you don't have enough expertise to evaluate the answer, yet that's where you most need the assistance, the support. And again, and this is a little bit worrisome about the overall narrative around AI, the place where we can get around this is with software development because

Let's say that I'm trying to accomplish something in a language that I don't know how to code in. I can have the agent do work for me because probably in the end I can run the program and look at the results. And that's what mattered to me, is that I run the results and I see, and if I don't see what I want, then I can complain and we can iterate. That verification step that doesn't imply I have comprehensive knowledge, it just implies that I know what I want to see in the end, is so critical and I think this is a causal factor in models being so good at coding because it's like the ultimate verifiable domain for them. But as soon as we leave that and go even into something like the legal realm, where the requirements are strict, but they're not codified in code and they have ambiguity about them, this whole picture falls apart. And you then are back at what you just said, which is this awful kind of paradox is like...

Yeah, use AI, but in the end, unless you're expert enough to evaluate every single one of its responses, you might be in real trouble. I don't know how to get out of this because the verification step is like we go to trial, but this is very consequential. That's expensive. Yeah, that's funny. I mean, it does make me think a little bit about, you know, some of the types of errors that we're trying to avoid are factuality and, you know, there is...

a temptation to say, well, let's just throw more tokens at it. I'll have a critic model that evaluates everything that is generated by the primary model. But then you go back to my observation that these models tend to correlate in their responses as well. Yeah, it's super interesting. That's a good point. Yeah, for my picture, we want real diversity of perspectives. This is just like red teaming.

For humans, this is most successful when you have a really diverse team of people who think creatively and differently. And if every one of the members of that team is thinking in a homogeneous way, they miss all of the crucial things. Same exact issue. If all the code review agents are biased in the same way, they will miss exactly the same class of bugs, and then we're all sunk. Yeah, I don't know how you'd encourage this diversity in the ecosystem or probably, as you say, converging towards some kind of one model. But I think for my picture, we need diversity.

Yeah, we got to keep those open weights models going or something because they're the weird players in the space. For sure, for sure. So we've talked about efficiency, interpretability, tokenomics, fluency. You're involved in a lot of different research directions. Excellent, excellent. What's next for you? Where do you see, either where do you see this all going kind of externally, but also like where is your research going?

Yeah, this is great. And as I said before, we're trying to think in weird and creative ways about what the future could hold. And I encourage my students to do this. And they're smart. So they say, all right, Chris, I'll think along those lines. But what's your answer to this question? So I do have an answer. And it's really shooting for the moon here, which would be, what about the architectural innovation that would upend the whole story around the stack transformer and the way we need to do data center build out to even get incremental gains in performance?

That could be upended, and it would come from some very innovative thing around maybe recursive use of the building blocks that we've got. So architectures, we should think. And when people say, oh no, we don't need more architectures, the transformer is good enough, that's where we should push back as academics doing something more clever and more scrappy that could change the world. And the other one is thinking in the interp space much more about data.

That's just because I want to tell the true story of how we go from data to model capabilities, but it also checks a box for me on connecting interpretability to safety. It has been hard for me to connect those two things. We have found some ways to do it, but it's not a slam dunk as a narrative, even though it's the dominant narrative. But I will say that when we get into things like data poisoning from innocuous examples, This is probably a growing societal concern. There is evidence that with very few examples planted in a pre-training dataset, you can have a significant influence on the outlook and preferences and quirks of the final model. So can we detect those examples? What's the nature of those attacks? How well hidden could they be? What's the smallest number of examples? And why does it happen? These are all going to be very pressing questions.

And so, again, it's just a data-oriented question that's very alive for me in the current moment. On the architecture front, are there, is the research that you're seeing or doing that is, you know, as yet under the radar that you think is, you know, promising and or underappreciated? I think you had my student Julie Kalidi on.

And she is an advocate for bite level models, essentially tokenizer free models. I think that's a big part of the future. It's a critical thing if you want to have truly multilingual models that are also equitable in terms of how many tokens they charge us for getting back to that earlier theme. But also Julie's perspective is that this is speculative, but I think there's something to this that it's a kind of inference time scaling because you do more compute at test time because you have more tokens and therefore more opportunities to build on interesting things. So that could be a big part of the future and the other one would be recursive architectures, as I said. But if you want to go all the way out, you could think, why do we always assume we're going to do gradient-based learning? There are lots of alternatives to that and nobody is exploring them because everyone takes it as a truism. We're all in our very narrow row here without even really realizing it. Who knows what's outside in this garden? It's very risky as a research bet because

Only one in a thousand of these ideas will pay off. But what's the point of being an academic researcher if you're not going to take that kind of risk? That's what we're positioned to do. Well, Chris, thanks so much for jumping on and sharing a bit about what you're working on. It's very cool stuff. Thank you. What a wonderful conversation. It gave me lots of new things to think about. Awesome. Awesome. Thanks so much.

Delete this episode?

This removes the episode page and its saved audio from this library.