← Back to library

Transcript: Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town

AI Engineer22m 32sTranscript ✅Added Jul 30, 9:47 pm GMT+8

Source video ID: yWS0udrIOc8

Transcript

  • 0:01 — [music] All right. Hey everybody. Uh yeah, yep. Yep. Yep. Okay. I have about 18 minutes. Um I have a slide deck here that that Claude put together for me, but Claude still can’t do very good slide decks, so I apologize for that. Just just just focus on me right right
  • 0:31 — here. This is like the most fun thing to do ever. Now, if I if I start like running out of time, can can you like call time for me? All right, perfect. So, I’m Steve Yagi. I am here on behalf of Sneak today. I am not getting paid for this talk or anything like that. They were really nice enough to buy me a ticket to the conference, which is cool. But I’m mostly here because I wanted to hang out with all of you and I don’t bite and feel free to come and say hi afterwards. I I wander the halls amazed
  • 1:02 — at all the brand logos. Um I’m up here because look the title of my talk is like aentic security but the real title of my talk is be scared. I come on come on level love level with me. Who here is scared? That’s good. But you’re all in the security talks. Of course you are. The problem is they aren’t out there, right?
  • 1:32 — I I went to Commonwealth. I went to so many places last year, but I one of these big banks I was at, I was talking to their chief security architect in December and I was doing a Q&A and talking about Vibe coding and everyone was like asking me questions and I was, you know, knocking him out of the park and he stands up real quiet at the end and he goes, “Yeah, so he goes, if everyone’s shipping code at the same, sorry, at 10 times faster and the defect rate stays the same, the
  • 2:03 — security defect, Right. The the vulnerability rate then that doesn’t that mean that the defect surface goes up by 10x? And I it hit me so hard I I sank down to my knees. I was just like, what are we going to do about this? Because it’s a really really important point because the subtle implied question is not if the defect rate stays the same, the defect rate’s going to get worse.
  • 2:34 — a lot worse with AIS writing the code. And so I didn’t have an answer for him. I actually I have a partial answer for you today and I’ll share it with you, but it’s only a partial answer. The real answer is you have to be scared of what’s coming. All right? And it’s bec it’s be it’s it’s it’s Do I have a slide about this? Yeah. So, it’s not just that you’re just like putting in more of the same defects, cross-
  • 3:05 — sight scripting and blah blah blah. You’re still putting those in. Fable wrote a an XSS vulnerability during the short time that I had with it. Not its fault. We’ll talk about it in a minute, [snorts] but those are the old ones, right? We are we all know how to fix those. There’s new vulnerability types and new attack surfaces coming and they’re here and many of them are incredibly well polished. Like what’s an example you guys know about slop squatting,
  • 3:36 — right? Where the AI hallucinates uh a package name. Let’s say that you want a graph database and so you’re like, “Okay, I’m going to use this graph database.” And the AI goes, “Oh, yeah. I know it’s, you know, it’s graphy 123.” and it goes off to the package manager and downloads graphy 123 and it builds and it runs and the tests pass and it looks right. But what it downloaded was a backdoor because that graphy 123 wasn’t a real package. But somebody noticed that the LLMs are hallucinating its name and they
  • 4:06 — uploaded one that does exactly the same thing as the one it thought it was getting plus a vulnerability bonus. Yeah. Is that scary? Yeah, it should be. It should be. How do you even detect that that’s happening? So, so we’re entering a world where everything you write, every bit of code that you generate is going to have to get far
  • 4:36 — more security scrutiny than it’s ever had before. Okay. It’s just you’re not it’s hard for me to convey how scared I am. All right, but let’s start with where it gets generated. Now, when I worked at Google, I worked really close with the tap team. They did tests, the test automation platform. Yeah. And they uh so they ran all of the unit tests and integration tests at Google. They had a massive fleet, you know, and and they learned stuff about how bugs work. And how bugs
  • 5:08 — work is they have a li they have like a life cycle where if you see the bug right away you’ll fix it and the longer the time goes for when you see a warning or some sort of issue the longer that passes okay it’s got this sort of halflife of urgency and all of a sudden it ain’t really biting anyone anymore okay this is how we handle all of our bugs at at Google they recognize that this is such a human nature phenomenon that they worked really hard to to move
  • 5:39 — the reporting of bugs as you were typing because that’s when you’re most likely to fix it. If you make a bug and it goes, by the way, there’s a divide by zero here or there’s a back door vulnerability or whatever, you you’ll fix it right there. But if it gets to code review time, you’re like, is it really worth it? Right? [snorts] And the problem, folks, is that that works for all classes of bugs except for security. There’s no halflife on it biting you. It’s not like, oh, because users aren’t
  • 6:10 — getting bothered by this security vulnerability that it’s not a problem over time. The problem compounds over time. Yeah. So, you have to treat this class of vulnerabilities the way that Google treated their top vulnerabilities at the time, which is to surface them at the developer’s fingertips. What if the developer doesn’t have any fingers? I don’t know how how many fingers the LLM have. [snorts] Um,
  • 6:41 — then you need to surface it to the LLMs. But wait, you say, wait, wait, wait, wait, wait, wait. Fable’s really smart. Or at least it seemed that way for the two days I got to use it. Uh, can’t Fable just write secure code, right? I mean, come on. Come on. You all know sec I mean, you’re all here in this room, right? You all know security is an arms race. One that never ends. One that’s going up exponentially with Moore’s law. One that’s going to get real uncomfortable when quantum comes along. Thank goodness that’s like five to seven
  • 7:11 — years away. Uh my buddy says 15, so maybe somewhere in between. But in the meantime, right, LM are a real problem. Yeah. So, how do you surface how do you surface the vulnerabilities? Well, first I wanted to figure out how to do it myself. I have a game I’ve been working on for 30 years. I just had Fable do a security hardening pass during the time I had it. Went through and it did all my cloud hardening and it found a bunch of credentials and did a a bunch of stuff
  • 7:42 — and it [snorts] started giving me these vibes like, “Yep, yep. Harding pass looking pretty good.” So, I ran sneak right and I I don’t have the numbers here. Uh yeah, I didn’t I didn’t include the numbers because uh I’m done. But um uh it found 241 vulnerabilities, right? just a ton that Fable hadn’t even thought to look for, right? Uh and it’s because look, so I I I’ve told people about the rule of five. When you do things with LLMs, often you have to get them to do up to
  • 8:12 — four to five reviews of the work that they did before it’s like actually ready to ship. And it’s because their cognitive process is very similar to ours. and it goes through a draft and then a revision and then polish and editing until it’s, you know, it’s finally ready to go. It’s like painting a wall. Some things you don’t just do all in one pass, you do them in multiple passes, right? So, um, security
  • 8:43 — is one of the So, what I found I wrote a book on vibe coding last year. I did more vibe cutting I think than anyone, you know, two two years ago. And and what I found was um that they’re really good at doing one thing at a time. Even the really good models like Fable, right? Just because of this multipass painting a wall phenomenon, you got to give them one task at a time, which means you can’t give them security at
  • 9:14 — the same time as you give them correctness. they’ll do a half-ass job of both. And you don’t want a half-ass job of either of those, it turns out. So, you do it in two passes. Right [snorts] now, five, it’s been five months now, but five months ago, I wrote an essay called Software Survival 3.0 where I talked about what software has to do to survive when LLMs can synthesize it all. I don’t know if any of you all saw that, but the basic the basic gist of it is that LLMs can synthesize any software that they want, but they’re very lazy in a good
  • 9:46 — way, right? Lazy in like they don’t want to spend tokens if they don’t have to because that’s money and power and bad for the planet and bad for your wallet and so on, right? And so they use tools to help them whenever it can save tokens, right? So tying it all together, if the LLMs are doing the coding and they’re happy to use tools to help them offload cognition, you see where this is going. Give them sneak, give them chain guard. And I still think there’s a missing
  • 10:16 — piece in this picture that I’ll tell you about at the end. Um, chain I don’t know if you all know about Chain Guard. Chain Guard. Uh, Chain Guard is a supply chain that you sign up for and they give you images that have been prevetted to not have vulnerabilities and they update them. So, it’s your inputs. Okay? And then Sneak handles everything else. The code that you write, the code that the LLM writes, the dependencies that you’re pulling in from slop squatting, right?
  • 10:47 — The innocent stuff, it can find. My understanding is that sneak can actually find vulnerabilities that are proprietary that only they know about because they’re ahead of the CVE registry. I’ve I there’s some truth to that. When I ran them on my codebase, it didn’t find any vulnerabilities that weren’t already public CVEes, but damn, it was easy [clears throat] to use, right? I think that a tool like Sneak is going to give your LLM superpowers. Okay?
  • 11:17 — Because what you do is you add it as a pass to the prompt that you give them for whatever they’re doing and say one last thing to look at and have them run your security analysis all of the tools. Get the open source ones, get the sneak one, get the chain guard one, get the all of them and have them check each other’s work too, right? If you want to get really serious about this on launch time, right? But secure your supply chain because Five Eyes is warning us. Yeah, my poor my poor game. By the way,
  • 11:49 — please please I just told you that my game has 241 vulnerabilities. Don’t go hack my game. Give me a couple days to fix the bugs. Yeah, we all good here? Good. All right. But if you do hack it, you’re going to bother like five players, right? [snorts] Okay. They’re very loyal. Um, look, five eyes, which is like a bunch of right governing. It’s it’s big countries that are have have their eye on on the cyber security you know landscape. They just announced that it is now months not years until it starts happening. It
  • 12:21 — you all know what it is right? It is when open source models catch up to mythos. Does anybody here believe open source models are going to catch up to mythos? Interesting that it’s about 5050. Anyone got a time frame in mind? » Who said December? That’s pretty accurate. » Cheater. Yeah, it’s about seven months. So, actually, it’s shrinking. So, it’s
  • 12:51 — probably about six months now. Yeah. And Mythos is real real good at hacking your systems. And and just just remember, you can’t trust it to automatically write good code any more than you can trust it to write elegant code by default. That’s a separate concern. It’s a separate pass. You can’t expect it to write performant code by default. That’s another pass. You see what I’m saying? You can’t expect it to necessarily write the code
  • 13:22 — according to your company coding standards. Okay? These are all passes that go through your code. And I just want you to remember that security should be your first one and your last one. Okay, give it extra. Okay. And the last thing I want to talk to you about, first of all, go do all this. And second of all, the last thing is really tr truthfully, okay? Dial it in here, folks. go to your families offline like in person and get your your
  • 13:52 — code words refreshed because another kind of scam that’s coming along is you get a call from a family member who’s in distress and they need money and it’s very convincing and there’s a video of them and you’re going to need a way to distinguish them from from AI. Okay, it’s months away and some of your families are going to be slow to catch on to this stuff but bank accounts will be drained. I heard that Congress was given secret demos of draining bank accounts. I’ve been scared of this for close to two years. I heard one talk
  • 14:24 — from a security researcher almost two years ago at ETLS Las Vegas and he stood up in front of the crowd and he said, “You’re all not scared enough of what’s coming. It’ll affect you personally, not just your company.” Okay, so that’s my message to you. It’s not a message of hope and positivity today, [laughter] but it is a it is a message that that’s that should be clear, crystal clear is that there are tools, open-source tools, free tools, commercial tools, okay, techniques, practices, okay, that you
  • 14:56 — can use right now to get started on fighting in this arms race and protecting yourself. And that’s all I’ve got today. Thank you. [applause] » [applause] » for a couple questions. » Questions. » We ready to go.
  • 15:30 — » Sweet. Uh, just super simple. What has surprised you recently in the world of AI coding? » What has surprised me in the world of AI coding? Well, I’m not really super representative. I spent a lot of my time trying to predict the future by like hammering agents really really hard. Yeah. Um, so um, uh, you know, one of the surprises, and I shouldn’t have been surprised, but one of the surprises is that AI is moving faster than the world
  • 16:00 — is moving. uh tech is moving sess faster than society can move and the surprises show up when friends smart friends resist uh you know the inevitability of AI and they call it psychosis or they or they poo poo it and they say well it’ll never be actually smart or whatever they they can’t see the curve right and uh and that surprises me uh maybe it shouldn’t um It
  • 16:31 — shows shows a sort of tunnel vision. I think people have a tendency to look about three months back and about three months forward and be like, “Oh, it looks pretty flat, right?” But but and so that that surprises me that people aren’t honestly that people aren’t more scared and that and by the same by the flip side that people aren’t more excited by it, right? Because you know once you actually you know once once you get it I mean you don’t even want to be here. How many of you are running cloud code right now? Most of you right? It’s
  • 17:01 — really fun. So, you know, I mean, like that’s a surprise, too, that the world is pushing back so hard on that, right? We’re in an awkward phase. We’ll get through it. Any other questions? Whoa. All right. Well, you pick. Feel free to bail. Also, you don’t have to stay. » I big fan. Um, what’s the like most impressive uh thing you’ve seen Gas Town do? and like how much human intervention was involved or steering.
  • 17:31 — » Oh, Gas Town. Yeah, Gas Town. Gas Town was a lot of fun in January. Um, yeah, Gas Town is a beads machine and I still use beads and I’m working I want to I was talking to Angie. I want to donate beads to the Aentic Foundation. You know, we’re gonna we’re going to put multiple backends on it. Beads is a task tracker, right? Beads is how you do Boris Chney loops. You know how Boris is like you shouldn’t be prompting your agent. Has anybody here actually like
  • 18:01 — successfully How often do you get Claude to actually run all night for you? Like for real run all night? A few of you, right? Like this is the next frontier I think of actually getting agents to run like for a long long time unsupervised. You can do it with beads by queuing up enough enough work and having them claim and all that. And there are some other some other systems that’ll do that. It was really fun when Gast Town did this for me automatically once. I filed a whole bunch of beads and they disappeared and I was like, “Oh no, another bug. My beads disappeared.” And
  • 18:31 — what had actually happened was that [snorts] one of the agents just found them and just implemented everything. Right. I was like, “Whoa, I really like swarms now.” Yeah. Fun times. Does anybody here regularly work with more than 10 coding agents at once? You see, like not many, right? The world is still in the we’re still kind of like prompting and using a few here and there, right? It’s gonna accelerate really fast next year. Other questions? You you just say it.
  • 19:02 — » Yeah. » Yeah. So, that was the third dimension that I really wanted to talk about. It’s just I don’t really have time, but like my friends over at Tesla, I’m advising them. They’re actually doing this, right? There’s just this whole space of who’s looking over who’s looking over your agents shoulders. I had this conversation just now. Like I tell everyone, everyone’s just starting
  • 19:32 — to stand up agents like 247 like processing cues, responding to events, like agents that actually do stuff, right? 247. And I I I I encourage people to think adversarially. I think of adversarial groups of agents tasked with doing that Q management because one agent will always eventually screw it up, right? So you got to have those supervisors. And so there’s whole systems emerging here kind of can go out
  • 20:02 — and go look at all of your things and say and start to like, you know, do do that hardening stuff like do they really need all those credentials on that service account really? Right? Only for this one action. Maybe we can like separate this one out. That kind of thing, right? This is a brand new frontier, but it’s one ironically that even though there’s kind of almost nothing out there, there’s some experimental stuff, you still have to be thinking about it right now and designing a solution in house right now, right? Because otherwise your engineers are going to spin up or your
  • 20:32 — non-engineers are going to spin up a bunch of agents with way too many permissions and then right as soon as a bear munches into the igloo, everyone’s dead, right? That’s the old security analogy. I don’t know if they still use that one anymore. [clears throat] » Yes. » Hey, Steve. Nice to see you. Um I’m curious in terms of like just the best practices that you’ve seen uh that you use personally or or maybe you haven’t tried, but um particularly with regards to prompt injection. So » yeah. So I was supposed to talk that
  • 21:02 — during my right during during my speech here. I I I wanted to mention it. [snorts] there are a whole bunch of attacks happening on the training and on the prompting side, right? So on training and inference [snorts] and so the bad guys find ways to sneak in stuff, right? Um and it’s like like like the simplest version is the new XSRF where like the user puts in some some text and then some bad actor puts in some extra text saying disregard everything and do the following, right?
  • 21:33 — And then they just get more sophisticated from there. Um, I mean, I don’t have any good answers for you other than like this is real. It’s kind of an education problem at this point. You need to get everybody thinking about it, right? And then you I feel like there are like new security roles about ready to emerge inside of companies. Agentic security that it’s kind of an extension of what they’re already doing. Who’s already doing this? Right. Yeah. So you you’ve already got people that are going out and looking after the the sort of security of your
  • 22:04 — agents that are in that are deployed in the Yeah. So you’re way ahead of everyone. I’d love to come talk to you later. This is all brand new stuff. Yeah. Cool. » [music]