Village Global Podcast - How One Hacked Library Can Take Down Thousands of Companies _ Feross Aboukhadijeh _Socket_
Summary
本期 Village Global 播客访谈了 Socket 创始人 Feros Aboukhadijeh,他是一位知名的开源开发者和安全专家,Socket 每周检测超过一千起攻击,保护包括 OpenAI、Anthropic、Vercel 在内的近三万家组织。他指出软件开发方式已发生根本变化:如今一个应用往往依赖成千上万个开源包,需要信任的个人维护者数量激增,使供应链成为攻击者最容易下手的软目标。他详细讲述了 Axios 被攻破的案例——攻击者用数周时间做社会工程,冒充合法公司,在伪造的 Microsoft Teams 通话中诱使维护者安装恶意文件,盗取令牌并劫持其 NPM 账户,短短三小时就波及海量下载。他强调 AI 编码代理比人类平均多引入 50% 的开源依赖,还会“幻想”出不存在的包名,被攻击者用来抢注仿冒(typosquat)。传统只看已知漏洞(CVE)的安全模型已不够,真正可靠的信号是实际阅读并分析代码的行为,而这正是前沿大模型让 Socket 得以大规模实现的能力。他还给出实用防御建议,如设置最小发布时间、使用锁文件、部署 Socket Firewall,并反思了从工程师到 CEO、从开发者主导增长到企业销售的转变。总体而言,他希望在 AI 时代守护开源的未来,让开源不因这些攻击而衰亡。
Highlights
-
I read the manual and found that there's this child lock feature. And so it was a pretty easy sequence of buttons. And I would just use it whenever I upset at my mom. I just sort of locked her out of the microwave and she'd get so upset and didn't know how to undo it.
我读了说明书,发现有一个儿童锁功能。那只是一串很简单的按键组合。每当我生我妈的气时,我就用它把她锁在微波炉外面,她就特别抓狂,还不知道怎么解开。
Charming childhood origin story of his security career -
During the call, the call just cuts out and a message shows up in the interface that says, your Microsoft Teams is out of date. You need to click this file to update it. And it turns out that file was not legitimate. And actually the URL he was on for the Microsoft Teams call its ...
通话进行到一半突然中断,界面弹出一条消息:你的 Microsoft Teams 版本过旧,需要点击这个文件来更新。结果那个文件根本不合法。而且他所在的 Microsoft Teams 通话 URL 本身就是伪造的。
Vivid, movie-like account of a real supply-chain social-engineering attack -
We do still see models making up package names that don't exist. It's almost like hopeful or wishful thinking where they say, hey, this problem I'm trying to solve, it would be a lot easier if this package existed. I'm just going to pretend it exists and let's try to go install i ...
我们仍然看到模型编造出并不存在的包名。这几乎像是一厢情愿的想法:它会说,我要解决的这个问题,如果有这么个包就简单多了,那我就假装它存在,试着去安装。当然,攻击者已经意识到,他们现在可以去抢注这些名字。
Surprising AI-era attack vector: hallucinated package names exploited via typosquatting -
We saw this malicious skill that was called what would Elon do and a lot of people were installing it to sort of inject a little bit of that Elon energy into their AI responses. And it became the number one skill because it turns out the skills marketplace that was hosting it jus ...
我们看到一个叫“What Would Elon Do”的恶意技能,很多人安装它,想给自己的 AI 回复注入一点马斯克式的能量。它成了排名第一的技能,因为托管它的技能市场存在一个极易被利用的下载漏洞,刷高下载量非常容易。
Memorable, funny example of how popularity metrics can be gamed -
The thing about security that's challenging for defenders is that the defenders have to be perfect. They have to guard every single way into the organization. Attackers just have to be right once. They just have to find one weak spot and they can get in. So it's a very asymmetric ...
安全领域对防御方最大的挑战在于:防御方必须做到完美,必须守住进入组织的每一条路径。而攻击者只需成功一次,只要找到一个弱点就能攻进来。所以这是一个极不对称的问题空间。
Crisp articulation of the fundamental asymmetry between attackers and defenders
Full transcript
He joins the Slack channel. The attackers were pretending to be a legitimate company. So they had sort of stolen the brand and the name of a real company. They got them on a Microsoft Teams call. During the call, the call just cuts out and a message shows up in the interface that says, your Microsoft Teams is out of date. You need to click this file to update it. And it turns out that file was not legitimate. Why we started building the company was we wanted to make sure that we could really help people understand what is the full scope the open source software that we're depending on and all the parts of that supply chain. And then how do we make good decisions about whether to trust it or not? And how do we be more proactive at responding to these threats? Hello, and welcome to the Village Global podcast. I'm Ann Duane, and today we're pleased to welcome Furose Abukadej, who's a security expert and an open source developer. He's very renowned, who started socket in 2021 to support the tidal wave of.
supply chain attacks that are happening now in software. Socket detects over a thousand attacks every week and they protect over 20,000 organizations from OpenAI, to Anthropic, to Vercell, and many startups. So without further ado, welcome for us. Welcome everyone and welcome Feros Abukadej and First, congratulations for us on your Series C at a billion-dollar valuation. Thank you. I appreciate it. It's a really exciting time for the whole socket team right now. Yes, exactly. Okay, so as a background, you're a serial founder and yourself a prolific open source creator and maintainer. And in fact, I think one of your claims to fame is your packages are downloaded a billion times.
A month? Is that right? Wow, so you're an open source billionaire. That's pretty good. We also have a lot of others on the socket team as well, like on our engineering team that are similarly prolific open source maintainers.
Amazing. Well, thank you, the unsung heroes of the Internet in the open-source community. But I actually would love to understand where your cybersecurity career began. And from what I've heard, it began as a child with a microwave. Can you tell that story? Sure thing, yeah. I did not know that this story was going to come up in this interview, so good job.
You did really deep research. Yeah. So the story is just when I was, I think I was like five. mom and dad got a new microwave. And one of the things I had a habit of doing then was I would read everything I could get my hands on, especially if it was technical. So I'd always read instruction manuals. There wasn't internet or much to read. So I would just read the manuals from front to back. And so one of the things I learned about this microwave was, I can't read the exact age now. Don't quote me on being age five. But yeah, I read the manual and found that there's this child lock feature. And so it was a pretty easy sequence of buttons. And I would just use it whenever I
upset at my mom. I just sort of locked her out of the microwave and she'd get so upset and didn't know how to undo it. Well, holding them hostage for their hot food. Well, I'm glad now that you use your smarts for good and not evil. And let's go back to 2020 when you started socket. And what did you see then?
Yeah, so before starting Socket, like you said, I spent a long time in open source. And one of the things that I started seeing more and more of was just how the way we write software had changed. And so we went from a world where applications had maybe a few hundred open source dependencies total to a world where you can't even get Hello World to show up on the screen without installing 1,000 plus dependencies. And this isn't a joke, right? This is people who build.
JavaScript apps today know this is the reality. And so we've kind of gone from this world where you had maybe a handful of projects you use that had teams behind them, foundations, this lot of support infrastructure to a world where you have individuals like myself who have individual authors of hundreds of packages. And so that meant that the amount of people you have to trust as you're building your software has just ballooned dramatically.
And so that led to this risk where there's now so many ways to get into a company because you can just compromise any one of these individual maintainers. And again, they're individuals. They're not necessarily organizations that have a lot of, they don't have necessarily all the infrastructure that a Linux foundation or a more mature project has. And so that's just become this softer target.
that attackers have recognized. And that's what attackers always do. They're always looking for what's the easiest way in. And that means that it's always changing over time, where they put their attention and where they decide to spend time. And my opinion at the time in the beginning of Socket in 2020 was that this was going to become an important problem. It, to me, felt so obvious that this is such an easy way into organizations. And I was actually surprised that nobody had recognized it sooner.
why we started building the company was we wanted to make sure that we could really help people understand what is the full scope of the open-source software that we're depending on and all the parts of that supply chain, and then how do we make good decisions about whether to trust it or not, and how do we be more proactive at responding to these threats. Yeah, very prescient, because it feels like a lot of people talk about code review of their colleagues who are carefully vetted for their own team, but again, they're accepting...
stuff developed by randoms on the Internet all the time. The order of magnitude is huge. I think I've heard you say that some big companies have 500,000 or more open source packages that they're running all the time. Yeah, we see up into the hundreds of thousands quite frequently, actually. There's a whole bunch of things that go into that, but there's also this version skew where you might be using a library, but you might have 70 different versions of it used by different teams.
the amount of different artifacts that you're depending on has gotten very large. And let's talk a little bit about the most devious, egregious things you've seen by bad actors in the ecosystem. And this is also where we scare the bejesus out of everybody and you get a lot of new customers. But maybe Axios or some of the other ones that you've been just stunned at how devious they were.
have been paying attention over the last three to six months. The first thing I want to emphasize is just that amount of these attacks has gone. It's like an order of magnitude more than we've ever seen before. And so just as one anecdote, on our end of quarter last month, we were busy. We're a busy team trying to close all these deals and do all this stuff. And we actually detected three major supply chain attacks that we had to write research posts about on that day.
You know, just this week on Monday, there was a huge supply chain attack that we also detected that affected Red Hat. And so it's like, you can't even go over a couple of days without one of these things happening. So the first thing is just the scale is kind of itself scary. But if you want to kind of zoom in on a couple of specific incidents to talk about. So this is a library that I think a lot of the founders on this call have used before. It's a simple library that lets you make HTTP requests.
kind of makes it a little bit more ergonomic and easier to do. And so you see it in a lot of applications. You see a lot of AI agents and coding tools also reaching for this library because it's sort of like the de facto standard in the JavaScript world. And so there's a maintainer who runs that project, and he received an email from a...
company saying they wanted to collaborate with him in some manner and invited him to join a Slack channel. By the way, I also received that same phishing hook myself. So they were targeting basically anyone who had...
A big following or a decent amount of open source code that was depended upon. And several other of our engineers also received that same lure. I think it's funny, by the way, that one of the best defenses against being phished is just being too busy to read your email. Because I didn't even see the email until after this attack had happened. And then we searched, and we realized, oh my goodness, a bunch of us also had that same lure, but didn't fall for it. But I can't take credit for falling for it. I didn't even see it, to be honest with you. But anyway.
So he got this Slack invitation. He joins the Slack channel. The attackers were pretending to be a legitimate company. So they had sort of stolen the brand and the name of a real company. They collaborated with him for about, I think, a number of weeks. I think it was almost a month of just sort of back and forth talking about different ways they could help his project and doing things together. They got him on a Microsoft Teams call. And on one of these Microsoft Teams calls.
During the call, the call just cuts out and a message shows up in the interface that says, your Microsoft Teams is out of date. You need to click this file to update it. And it turns out that file was not legitimate. And actually the URL he was on for the Microsoft Teams call itself was not a legitimate URL, but everything about it looked.
exactly the same because the attackers use the official Microsoft team SDK to make like all the components and the buttons and everything look identical. But, but of course, like they were able to kind of code in this behavior where at some point in the call, it would stop the call and then ask him to install this file. And so of course, you know, we've all been there. We've all been like late to a meeting or like, you know, you're not having technical difficulties and we're like, come on, I just want to get back in the call. What do I do here? You're scrambling, right? And so that's like the moment where your most
vulnerable, right? And then the thing is, people have their guards up usually when they're first interacting with people, but not after our relationship has been developed for that many weeks. And so it was a very effective social engineering attack to do. And then just to tie up the whole story, that file that he ran ended up being...
a malware that gave them access to his entire system. And what they did was they pulled some of the tokens off of his machine that gave the attackers access to publish his NPM packages. So they were able to basically take over his maintainer account and, therefore, put their back door, their malicious code into his project, and then, therefore, propagates downstream to everybody who's using that library. And again, the thing about Axios is it's so widespread.
At the moment, you compromise that. I mean, it has 100 million downloads per, I want to say per week, just a wild amount of reach. And so even if you compromise that for just a few hours, which is how long it was compromised for, it was a total of three hours. The amount of downloads that happened in that period and the amount of agents that decide, hey, we're going to NPM install this. During that window, it's very widespread. And so they were able to steal a lot of credentials from the people that ran that package during that three-hour window.
Wow, crazy. And maybe let's talk about one or two more, and then we'll talk about how Socket helps. Maybe how about the Mercur Breach that got impacted OpenAI and Anthropic and others? Yeah, so the Mercur Breach was a similar story to Axios. It was a library called Light LLM, which is used to do LLM routing. And it's a very widely used library.
almost identical kind of story. One thing I will say that is interesting is that these attackers, because they're attacking through the supply chain in this way and they are running their code on developer systems, they often, when they steal credentials, they're getting a lot of other maintainer accounts and it becomes almost worm-like where...
you can spread virally through the network by taking over one NPM package. That gives you access to everyone who installs that. And then we see these incidents where 700 or 1,000 packages are compromised in a few minute window. And so that's really scary if you have developers that publish open source. It's not just that your own.
Your own infrastructure could be affected. Your own apps could be your data and things like that could be affected. But your open source libraries and SDKs could be themselves used to spread and attack other organizations because you have a developer on your team that just has a token sitting there on their laptop.
Yeah, scary. OK, so let's talk about how Socket works and then how you get it. Because as I understand, you have a free tier. And it's GitHub marketplace. Click, click. You're in. But tell us a little bit how it works. Sure thing. Yeah, so the easiest way to get started is to go to GitHub and, sorry, go to our website and there's like a two-click install. And Socket will connect into your GitHub. And we'll start looking at every commit that any of your developers make or your agents are making. And same thing with pull requests.
So really, anywhere where new code is being brought into your application, we want to take a look at that code. And we want to make sure that we're assessing it. Got it. So you're doing the work of trying to find if there's bad actors in the packages that a developer is using. And so you're doing the work they don't have to use. And that applies to both new installs and then also any updates. Correct. Yeah. Updates are also, you know, as we talked about, you could be using a trusted library, but then it turns out the next version of that library has been, you know, it's been compromised. And so we really need to assess any time new code is coming into our organization through any method, whether it's a developer, whether it's an agent making the decision, it doesn't matter. That code, it needs, we need to have some
some process, some trust boundary where we say, this is not our code and we need to know that it's safe. So Socket is a really easy way to get all of our intelligence, all of the threat research that we're doing and just get that kind of automatically available to your team. For everybody on the team. And it's also important I want to add in to protect the endpoint too because we don't want to wait until...
a code is pushed to GitHub to learn that it's unsafe as we talked about. It's actually been one of the more interesting trends over the last year or so is how important the endpoint has become. Because for a long time, there was this conversation about how we're going to move to doing everything in cloud sandboxes. And the developer workstation doesn't really matter anymore because we can just use all these cloud resources to do our development. But with things like cloud code and even co-work for the non-developers on our teams, those operate directly on the device. And so it's been this shift back to the importance of our actual devices. And there's been this race to adopt MCP servers. And there's a lot of access that we're giving to these agents. There's a lot of access where we're giving to the MCP servers. We're putting tokens into...
files on our systems so that the MCPs can access a lot of things. And so when an attacker gets onto a developer device or onto anyone's laptop, it's a very rich place to kind of find scary. There's a lot of tokens and things that are on there. And so we actually need to protect the device too. And so last October released something called Socket Firewall.
And this is a way to protect the endpoint. So the way to think about it is, it doesn't matter what agent you're using, or what harness you're using, or what tool you're using, any time third-party code is being requested by your device, we want to assess that. And we want to block it if it's known to be malicious or not to be safe. And so that's a really easy thing for folks to use, also free as well.
Great. We have very generous free time. Awesome. Awesome. Great. Okay, so let's talk about that, because since you started in 2020, a few things have happened in the software development world, including coding agents. And you've said that coding agents, on average, pull 50% more open source software than humans. First, why is that, do you think? And then what should companies be doing about that?
An agent will often reach for dependency because it's the fastest path to solving a problem. And they're also trained on huge corpuses of open source code. And so there's a lot of open source code that informs how agents write their own code. And so in this world of like I was talking about earlier in this new world where there's just dependencies flying around everywhere and our apps are built on this mountain of open source code, it's not really a surprise that you have agents.
basically building software in the same manner as the human code on which they were trained. There's also some risks that come with that. So obviously, we all know about hallucinations. I think that's gotten a lot better as models have gotten more powerful. But we do still see models making up package names that don't exist. It's almost like hopeful or wishful thinking where they say, hey, this problem I'm trying to solve, it would be a lot easier if.
This package existed. I'm just going to pretend it exists and let's try to go install it. And of course, tackers have realized that they can now go and typo-squat the names and hope that agents will come and actually install those. And so this is why it's really important to change the traditional security model we've had where In the past, people really focused on what are called known vulnerabilities, or CVEs. Sometimes you'll hear them referred to as. And these are the typical software vulnerabilities, or basically bugs or accidents in software. But what we really need to do is think about a much more expansive definition of trust. Because when an agent is looking to install something, and that package, let's say, has
50 downloads like no one's using it effectively. It's a very very niche Maybe it was published five days ago, right? No one no one has heard of it in the traditional security scanner world before socket Tools that the companies would use would scan that code and say hey There's no known vulnerabilities like I haven't found any reports saying that this is vulnerable But the reason for that is nobody's bothered to look at this code if no one is using it, right? It's literally completely random code. But that doesn't mean it's safe, just because it has no known issues. There's probably a lot of security vulnerabilities in it, not to mention the entire package itself could be malicious or could be designed as bait for your AI agent to install. And so that was a big part of the shift when we started the company, was we wanted to make sure that people could, from first principles, take an untrusted piece of code and make a determination about whether they want to use it in their application or not.
And I think there's such an important point. And Pierre, one of the founders at Tesorio, asked, GitHub stars can now be kind of bought via marketing. And sometimes popularity can be also bought via marketing and other things that are spoofed. What are the real signals that either you're looking for or developers should look for on these? Yeah. Well, I think that the source of truth at the end of the day is what does this code do when we run it? At the end of the day, open source code is going to be bundled into your application that you're shipping to your users, or that you're pushing into production. And the operating system doesn't care whether you wrote the code or whether an attacker wrote the code or whether an open source maintainer wrote the code. It's all part of your code at the end of the day.
people need to sort of have a mindset shift around this because for a long time people thought well it's open source it's not my problem right I'm sure the community will take care of it or it must be good because others are using it right but we've seen with things like Axios it doesn't matter how many people are using it hundreds and millions people could be using it and it could still get compromised right and so You know superficial metrics like stars even download counts can be faked right we saw this malicious skill that was called what would Elon do and a lot of people were installing it to sort of inject a little bit of that Elon energy into their AI responses and It became the number one skill because it turns out the skills marketplace that was hosting it I'm not gonna call it which one it was but they just had a trivially
exploitable kind of download. It was very easy to inflate the download counts. And so they became the number one skill, even though they didn't really have any organic usage. And of course, then they got a whole bunch of people installing it because it was number one on the leaderboard. And so these metrics are often gamable. We've built some detections to figure out when people are buying stars, when we see like a natural star networks and things like that. But to answer your question, what should people look for? You need to have some something that is ungameable and tied to ultimately what is the code going to do? And so I'll shout out socket because what we do there is we look at the code and we figure out, does this code access the network? Does it read your file system? Does it read your API keys and your environment variables? What is the code doing? That really tells you whether or not you should trust it. All the other things around it, they can be useful indicators, but ultimately,
there's no replacement for just reading the code. And obviously, humans can't read the code because of the scale we're operating at, but we can have AIs like Sockets AI look at every line of code and ensure that nothing is getting in that hasn't been vetted, right? Right. And again, you're doing the work so that everyone can benefit from that. OK. And we can afford, by the way, to do quite a big investment into the work we're doing, because once we assess a package, that's often used by Large numbers of our customers and so we can afford to put an unreasonable amount of tokens into really trying to understand and characterize Whether one of these things is trustworthy before customer even requests it great. Okay, let's Move to vibe coding so what is the process you recommend for an organization to put between a vibe coder who may not be technical at all and
the environment. Yeah that's a really great question because we're seeing so many non-technical folks now, vibe coding and building apps and it's awesome. I love to see it. As someone who taught computer science for a number of years I really love to see more people.
coding, and I think it's going to help more people get into the field and just empower people. It's awesome. That said, like we've been talking about, developers haven't necessarily done a very good job of assessing risk of third-party code. And so how are we supposed to expect someone in sales or marketing no offense to those folks, but to understand what this code is doing? And so we need guardrails that are in place to sort of let them loose and to let them have the power of all these tools without having any of the downsides. And so the simple solution is something that we've, again, I feel like I'm just an ad for socket. We've been thinking about these problems. And so I do think we have a lot of the best solutions here. There's something called socket firewall that I mentioned a moment ago that you can put it in place. And it will make sure that nothing the agent does.
as it relates to third-party code is unsafe. So the agent can't go out and fetch this untrusted package that we've been talking about. It can't go and install this random library that nobody should be using. It's going to guard against that. And that's really powerful because it means you also can let your teams experiment with whatever AI tools they want because we're doing this at the network level. We don't need to be hooked into every new AI tool or have like a...
You don't need to prove what tools your team is using. You can kind of put this one control in place. And then it doesn't matter what AI they're using. It's all ultimately going to transit the firewall. And we're going to be able to see it and act on it. Great. OK, great. Another new thing since you started is LLMs. And Zack Stone at Google had asked a little bit about, why haven't we seen more bad stuff via prompt injection? And is that something you think about?
Yeah, I think attackers always look for the easiest route into an organization. And they're really great at hunting for soft targets. This is one of the things, by the way, about security that makes it fun, is that it's a cat and mouse game where there's attackers, there's defenders. And when an attacker improves, a defender can respond. And when a defender improves defenses, the attackers respond and go somewhere else where the defenses are lighter. And so what I think what we're seeing with prompt injection is that it's just not the easiest way into organizations right now, because it's sort of finicky to get them to work. It's often unclear how to get data out when you've done a prompt injection. And I think things like the supply chain have just become easier. Yeah, why prompt inject one organization when you can take over Axios and have tens of millions of installs? The attackers that did Axios, they're called TeamPCP is the name of the group. And they claim that
from, I don't think it was another one of their compromises they did earlier, but they claimed just from that one compromise, they were able to steal 300 gigabytes of credentials. And those are compressed credentials. So imagine a text file, compressed 300 gigabytes, you're talking an untold number of organizations tokens and keys. And so when you have things like that, I don't think we're going to...
It's going to be like this forever because, as I mentioned, it's really an evolving system. But I just don't think right now that's where the easy opportunities are. Got it. And to that end, if it happens and either socket notifies someone or they hear that they have had some kind of supply chain attack, what should they do? So the best thing is to not be affected in the first place if you can. So definitely get these tools in place. We already talked about.
socket for GitHub and the socket firewall. There's other things you can do as well. I just want to mention to make sure for completeness people know about real quick. So people should definitely go and set up a minimum release age in their package manager. This tells your package manager to not bring in code that's newer than a certain number of days. So you can say, what's the threshold would you recommend? I would say, you know, at least 24 hours would help a lot. A lot of the things that socket catches, we catch them in a few minutes. And then we work with GitHub and we work with npm to get that code taken down to protect.
entire community, whether or not they're socket customers. So if you set your minimum release age to 24 hours or to three days or something like that, you will downstream get the benefits of the work that teams like socket are doing, even if you're not a customer. But ultimately, that's a pretty good defense, but it's not perfect. There's malware that slips through that lasts longer than three days. The other thing that's interesting is there's a tension between When you set that, let's say you set it to seven days, right? You want to be really safe and you want to use any code newer than seven days. Well, now what do you do when there's a vulnerability? And obviously, as we know, I mean, mythos is finding vulnerabilities in a lot of open source code. The other frontier models are doing the same. And so we do need a method to bypass that waiting window and actually update our code to stay safe from vulnerabilities. And so it's this.
tension, direct tension between, as an industry, we've been trying to get people to update faster and faster. We always say, update your iPhone. Don't run old software. You're going to get attacked. You're going to be vulnerable. And so after a decade of trying to get the industry to take this seriously, we've gotten people to realize it's very important to update just in time for supply chain attacks to come around. And the faster you update, the more you're running code that nobody's vetted, nobody's looked at. You don't necessarily want to be running code that was published four hours ago that no one's had a chance to look at.
And the malicious actors are bad. Sometimes they can wait, like, hang out for 24 hours, right? And then they release their encoded action or whatever. And let's just talk about that. So many of the folks on the call probably are working in organizations where the CEO is saying, lock it down. You shouldn't be, you know, we got to get serious about this.
But open source is really important. So how could you help them navigate those waters? Well, I mean, I'm glad to hear there's CEOs that are saying, lock it down. I've heard it's usually the opposite. We've got to do AI everything, AI everything, right? And pushing people to kind of...
just go for it and ignore the costs. If there's CEOs saying that, that's good news to hear. I definitely think we're seeing more of that. It's definitely become a board-level concern in a lot of the conversations we're having where boards are saying, why were we affected by this? What are you going to do to make sure we're not affected by these kinds of things in the future? Or one of their other portfolio companies was affected, and they want to make sure that everybody is protected.
focus on that even from some CEOs. Got it. So we talked a little bit about what startups can do, or individual developers. In a large organization, maybe that's been around for a couple years, they were pretty optimized for a SaaS kind of environment. Is there any suggestions you'd have of what they should start, stop, or continue doing besides using Socket and other things? I think the most important thing is, and again, we're talking about supply chain here, since that's the topic. There's obviously a lot of other aspects to security.
That would be a longer conversation. But when it relates to supply chain, the number one thing is you need to understand what open source code you're using. What is part of your supply chain? And by the way, supply chain is actually more expansive than just open source. Right now, open source is the one we're talking about. But even things like the APIs that you use, the various vendors that provide you software, these are all dependencies that Effectively, it's other people's code, but it's your problem. It doesn't really matter if it's their fault, if ultimately you're the one who has to explain to your customers why you were breached or why you had an effect. So you have to think of, really, what is all the code that could affect me? And you need to have an inventory is what I call it. You need to know what that code is. And so step one, figure out what you're using.
something socket can help you with. But as a baseline, you need to make sure that all your projects are using a lock file. This is the default in pretty much all package managers today. But what a lock file will let you do is it locks down what versions of open source code your application is using. If you don't have a lock file, you're sort of uh, getting whatever happens to be like the newest thing that's out there or there's usually like these loose version ranges where you're saying, I'm okay with anything that's, you know, version 1.2 point, whatever. And so that means that when something is compromised, the next time one of your developers runs an install or does a build, you're just getting that new code. And there's no way for you to know, like it's very time dependent almost like you might get unlucky and do a deploy, you know, when something is malicious or vulnerable. So you want to be very intentional and a lock file is just like baseline and everybody should be doing this. If they're not like.
Go do that today. It's not hard. That now creates a definitive inventory of like what you're using. Um, and then, um, you need some type of guardrail using that inventory to say, if anyone's changing that inventory and bringing in new code, we need to be able to decide whether we want to allow it or not. And so there's many ways to do that. Obviously again, socket has a way to do that, but, um, but you need to know, it needs to be an explicit decision, right? It needs to be some type of, um, tool informing whether or not you want to do that. I think that's just step one. And you'd be surprised the number of companies that still don't have lock files. They don't have these kind of basic things in place. Great. Well, let's pivot a little bit and talk about your business. So you're a developer led.
growth company, but you also are now doing a lot of enterprise sales. And I think there's 20,000 organizations, including Vercell and OpenAI. Almost 30,000. 30,000? Okay, I can't even keep up. It's definitely inflecting. Can you tell us a little bit about...
what you've learned about developer-led growth and then the move to enterprise? Yeah. So when we launched Socket, we were just a GitHub app. There was no website. There was no dashboard that you logged into. And it was literally designed to be something that a developer could just install.
And so we've always had this ethos of we just want to put the tools in the developer hands. There was no billing or pricing or contact sales or book a demo, none of this stuff. And I think that came from my and the original team's background as open source maintainers, where we really care a lot about making things that developers want to use. And so that's what you do as an open source person is you make things for other developers.
where we how we built the product. And that also showed up in the way that we provide our data. So you can look up any open source package on socket and we give you everything we know about it. And this is our crown jewels, really. This is like what we spend an inordinate amount of tokens and compute to analyze and to determine. And this is what we sell to our enterprises. But we also provide it for free on our website. And this is great because it means people can come in.
search for their favorite libraries and see, okay, what does Socket know about it? Does it find anything interesting? You can click through and see the recent malware that we've detected. That makes it really concrete for developers because a lot of times security can be abstract. It's like this Hollywood movie kind of scary stuff, but if you can actually click in and see, oh my goodness, this is a library that...
I use, like Axios for example, and you can literally see the code of the malicious version and see, oh my goodness, look at what, I can read the code and see, this is what the attacker is trying to do, they're trying to steal this and this and this and this, and you go, oh my goodness, I'm running these commands every day, and who knows if one of the times I do that I'm gonna bring in some of this code and they can actually see examples of it, right? So that was very helpful to us in the early days, to just teach people about this problem that nobody, I would say not nobody, but very few people had as a top of mind problem. And so we were able to raise awareness and really educate people on it. And for a long time, the free version of Socket was all we had. We didn't have any type of billing or upsell or enterprise plan. And then of course, with our first customers, the enterprise, the first enterprise sales, we're really just selling the free product with support and extra assistance to our design partners. And we sort of
built into an enterprise plan and started building newer and newer things that we only put into the paid plans. But we still kept that really generous free offering for our users. Yeah, that's incredible. And who is the buyer at the organization that you usually work with? Usually it's the security team. So it could be the CISO or the head of application security at a larger company. And more and more, though, we're also starting to see the folks that are in charge of AI rollouts, like AI coding tools, or innovation offices in bigger companies, they might be responsible for rolling AI out, but then doing it in a safe way and making sure that it doesn't compromise the company. And so they're very interested in how to put guardrails around the AIs, so they can let their employees run free.
Usually these amazing tools, but to do it in a way that's where the risk is contained. Got it. And what have you learned about moving from founder to CEO now with a sales organization? What have you learned there? Oh man, it's been a huge journey because I've been a developer my whole life. I've been an engineer and very much into the weeds of code and trying to write good code and almost like craftsmen of code. That's kind of how I always thought about myself.
find my job changing so much. As we've been doing the company, it's been really quite a journey. I like sales now, actually. So I came into this thinking, sales is bad. It's dirty. The product should speak for itself. I don't like it when I get cold calls from salespeople. I don't like it when I have to deal with salespeople. That's kind of how I came into it. But really, all sales is just like, if you have a good product, you should want people to use it.
you should help them understand why the problem is important and how you can help them. And I think there's a way to do sales that is very true to that engineer spirit, where you really are consultative and you're like, you'll tell them honestly where the product is good and where it's not. And it's really about fit and about being...
a consultant for them. I love it now because I just get to talk about the problem and my passion for the problem comes through when I'm talking to people. I think it's important to balance it with the engineering first culture, and you don't want to ever have one dominate the other. There are so many companies I've seen that have just amazing engineering, but they can't figure out how to get the world to care. I want to maximize impact of the work we're doing because I think it's very important. And I really want open source to have a bright future. And I don't want these attacks and these risks to make people, especially with AI agents coming out, and how easy it is to just write code. I don't want open source to die. I think that would be really bad for the world. And so this is a way for me to have that impact. But you don't want to, I've seen so many companies where they are
have technically excellent product, but they just can't figure out how to get anyone to care. And so that's a way to die, right? The other way to kind of die as a company is to go too far into the sales side and lose track of the fact that the product actually has to be good. And so I see that happen to some of our competitors. They had a great product, but then they...
put in a sales CEO that under-invested in engineering, and they just kind of went hard into sales, and that meant that they grew really, really fast. But then they hit a ceiling, and they've not been able to get above it, because now customers have realized the product hasn't improved in five years now. They've under-delivered versus expectations or something like that. I think they're both...
awesome. Sales is awesome, engineering is awesome, and marrying the two and being true to our engineering spirit, but also not being afraid to do sales and to actually go and win. It's been this nice balance that we've been able to strike at Socket. That was great. George Forman, who did the George Forman Grill, which is one of the best consumer products in the world, he always said, I don't know anything about sales, but I know how to talk passionately about a product I love. And that's kind of what you do. Pretty much, yeah. Well, let's I'll talk a little bit about socket engineering and development. It seems like your product has really benefited from the improvements in the frontier models. And can you talk a little bit about how you surf that wave and think about that? Yeah, so when we first built socket, our insight was we think there's going to be a lot more supply chain attacks. And the scanners that focus on the known vulnerabilities are just missing a big part of the picture because they're literally not looking at the code, right? Just to really drive the point home.
The way vulnerabilities work is there's this government run, federal government run database called the National Vulnerability Database. And it's kind of a clearinghouse for all known software defects. And people push those reports into there. And then all the vendors do is basically pull down from that database and tell you if you're using any software that's on the list. So it's a very simplistic kind of thing. And so we thought.
know, if we really actually analyze the code, we're going to find a lot of stuff that people don't know about today. And we're going to find the really interesting threats like these these nation state attackers, this malicious code, and that this is going to get worse. Right. So that was our insight. And the way we set about solving it was using our our experience as maintainers, we thought, if we see certain signals in the code, then we know that that's a sign that something about this package has changed that makes it riskier to use. And so the way that we thought about it was, say, you use iPhone or Android, whatever OS you use. But on your smartphone, when an app updates and it wants to use, let's say, new functionality, like your camera or your contacts or your location, it doesn't just get to use it because it wants to. It has to prompt you. The OS actually enforces that.
significant changes in risk have to be approved by the user. Unfortunately, open source and the way we code today doesn't have that type of permissions model or capability model. Our insight was if we could just find out when that type of change happens in code that we depend on, we would be able to spot a lot of these attacks. That was true, that worked. But what happened was it was too much data. It was actually overwhelming to most of our users. There were certain really technical teams that really loved to know every single time one of their libraries, oh, it's making a new network request here. That's really interesting. Why is it doing that? And to be honest, that's the kind of stuff that interested me.
It might not be a security issue, but I might tell you something interesting about, again, because it's all part of your application. Right, functionality or whatever you want. Why is it doing that? Is that going to affect our users? How does that affect the trustworthiness of this code? So it's all interesting stuff. But what we learned was only the really most technical teams who had really solved a lot of their other security problems and had time to spare to look at this stuff found that level of granularity useful. And so what really made socket.
useful to a broader set of people was AI models and frontier models because we were able to take this just amazing set of data that we were producing about libraries and put it into the models and have them synthesize it and really boil it down into something simple that anyone can understand. So really you go from all these signals, you go to really almost a Boolean.
Should I allow my developers to use this library or not? Yes or no? Really simple. And that's what people really wanted. And before models, there would have been no way for us to do the...
do this at the scale that we needed to do it at across every single library, except in real time. But with models, it's almost like you have this infinitely scalable army of interns. It's actually gotten better than interns now. I'd say they're like PhD level. But yeah. When we started, we started this with GPT-4. We got early beta access to GPT-4. And that's kind of where this first started to really work, because GPT-4 was just the first model that was just good enough to really start to.
to do this work at scale. It had a whole bunch of false positives. So we had to include humans in the loop to kind of approve things. But it was very, you could see that it was going to be a big deal. And yeah, and so it's like now it's obviously gotten way better and better. And it's just amazing now we can catch these attacks in like minutes within.
minutes of that code being out there, our system is kicked into gear, and it's just pulled apart the code, and it's found all the problems, and it can make the assessment really, really quickly. That's amazing. And you've made the point that it used to be that the attackers had all the incentive, and the defenders were like saying, hey, I have small resources. How do I defend against all these things? But now with agents, AI agents, and the LLMs, actually you can afford to defend more, right? And certainly Socket can afford to.
scan the environment but more. And that's absolutely true. And I think that the thing about security that's challenging for defenders is that the defenders have to be perfect. They have to guard every single way into the organization. Attackers just have to be right once. They just have to find one weak spot and they can get in. So it's a very asymmetric problem space. You have to be perfect as a defender or else You know, the attacker gets in. And I think AI, the big question people have been asking is, like, is AI going to help attackers more or is it going to help defenders more? And I think it really depends on the specific part of security that you're talking about. So for things like phishing, I think AI has helped attackers quite a lot more because they can write these amazing phishing lures that are customized to exactly who you are and everything you've ever written or done online and everything about your organization. There's a lot of, you know.
open source intelligence out there that they can use. But for defenders on the supply chain side, I think it's actually been a net benefit to defenders because you have this world where people were just pulling in code and never looking at it. And now you can actually have at least agents looking at it through things like socket or others. Great. And let's talk a little bit about a question we got relating to private information, like from a health care founder.
How would you think about partitioning that confidential information? Is there anything unique about that or anything, any advice you'd have? That's a very deep question. I don't know if I can give a super short answer about that, but I think if you read the healthcare space, I think you probably want to do your first security hires sooner rather than later. There's a lot more regulations around health data. You want to think very carefully about...
allowing models to access that data and giving them tools that can pull information on various customers. I think one of the things people need to keep in mind with putting an agent in front of anything like that is that that's where prompt injection really becomes a serious issue. I don't know if you saw recently Instagram rolled out a support bot that helped people get access to their accounts if they lost access.
This is important for them to do, because Instagram notoriously has really poor support around getting access to accounts. And so they were trying to roll this out to help burn through their support ticket queue a little bit faster. But it turns out it was trivial to convince the bot to just add.
any email address to anyone's account and then you could then go through the password reset flow and reset their account and get into their Instagram account. And so like the way to think about these things is like if the agent has access to a tool, you should just assume like all your prompts and your instructions that you're giving to the agent can be ignored or can be bypassed through a clever prompting by an attacker and just assume that those tools are directly accessible by your users and that the agent that's mediating it and trying to sort of be that human, you know, support.
like intelligence, just assume that that can be tricked. Just like you would for any human in your organization. You have to assume that even your own support agents at your company could be tricked or could be bribed inside your threat. There's a lot of similarities actually between the way we think about the AIs and the way we think about the risks of humans in our organizations as well. Amazing. And let's talk a little bit about people are always curious about how modern engineering organizations are constructed and organized. Is there any learnings for you about, you know, building a really AI-native company? That's a good question. We had one founder recently post. He's like, it's not a democracy anymore. It's like, you know, I, because a lot of people, I think what the sentence was, people were testing a lot of different tools and vibe coding and kind of a lost a little bit of the focus on, OK, what are the core metrics and things like that? Yeah, I think that there is there's
definitely a risk that people can go too far with the sort of AI native stuff. So we've been at Socket, we haven't been pushing it on people in any way. Some CEOs are telling everybody, you know, we got to go all in on this. We've sort of been very organic about it, like letting developers experiment with it at their own pace. And it's naturally become something that most people on the team are using. And I think that's at least for me, that's the way to go. I don't think doing anything unnatural. But yeah, I think so many thoughts on this. It's a very big time. Yeah, we'll do another podcast on engineering works. And we're coming up on time. So is there anything I didn't ask you about or anything that's top of mind for the future of socket that you want to share? I just think it's such an interesting time to be doing.
cybersecurity company and also just any company really. What a blessing it is that we get to live in this interesting time, in this interesting moment. It's such an exciting time. Every day is different and the pace is so fast. For all the founders that are on this call, there's so much opportunity.
really lucky to be part of all this. That's great. Well, and thank you for all you do to keep the companies safe and individuals safe. And we're going to post some links, not just to socket.dev, but also to some of the videos you've posted as you're discovering some of these supply chain attacks in real time and advising people what to do. And it's really fascinating to watch. So it's like reality television. Sometimes these attacks are just like right out of a movie. It's kind of fun to, people want to go and check out our blog at socket.dev slash blog.
some of these attack. I mean, there's, they're like every day, right? You can, you can see, you know, it's, it is really like a thousand a week you're detecting or some crazy number, right? Yeah. Yeah. Yeah. And we, not, not all of them are interesting enough to, to merit blog posts because we do see a lot of same type of thing over and over again. But the ones that we write about are just, they can be very interesting. And do you name them? Um, cause some of them are named after like the, the worm in the movie Dune and others. Do you get to name them or does somebody else name? Usually whoever finds them first, uh, can coin a name for them. And, and so yeah.
We've tried to, you know, come in fun names when we can. Yeah. Wonderful. That's part of like your, it's like, you know, as the discoverer, you get that privilege. Yes, like the stars. Right. Well, for us, thank you so much really for all you do and such a pleasure seeing you today. Yeah, thanks again. Appreciate it. Hey, this is Ben Keznoka, co-founder of Village Global. Thanks so much for tuning into the Village Global podcast, where we go deep on all of the biggest topics in tech. If you enjoyed this conversation, please subscribe to our YouTube channel. You can check us out on Spotify, Apple, wherever you get your podcasts. We'd love to see you for the next one.