AI Horror Stories (Ep. 164)
In this episode, Isaac Askew and Dustin Rea share real-world AI security stories, highlighting the importance of security best practices when deploying autonomous agents. From data leakage to infrastructure risks, learn how to protect your systems in an AI-driven world.
Episode transcript
Welcome to Never Write. I'm Isaac Askew. Dustin Ray. And I'm Jeffrey Sherman. And today, we're going to give you crazy AI stories. No, there's no giant overall thing we're just here to make today. It's just for the giggles. Uh so, Dustin, you want to kick us off? Let's give us a crazy AI story that's happened to you. Yeah, I'll kind of start off with I think one of the more common ones. So, like I have an an agency or you know, freelancing business I work with several clients at the same time. I have my own work in linear as well.
Of course, I'm connected to linear for different clients as well. There was a time and this is I guess a case for why I think human oversight is still important even in in autonomous agent workflows. I caught my agent pulling context from my business and including it like in a client's like context like to work on like a plan like it was including things that I just talked to it about in a different project. I don't know I can't remember exactly why in this case it had that access.
But that's what led me to the case of like it they all have to be like in isolated VMs now. Yeah, the the problem is the leakage of data. Was what it was talking about at least relevant like you were solving the same problem for multiple customers and you had Yeah, it was more or less relevant in that sense where like it was I can't remember exactly what it was but it on like some of the frameworks that I'm working on or like my starter kit or it was something like that that I I was working on.
It was like pulling that context into the client like to like use it already and I was like no like this isn't these aren't connected things. Like these are there should be a wall between these two things. Like that's you know, and I hadn't I hadn't ran into it until I saw it happen you know, personally and then I was like okay, this is something we got to you know, take really seriously. I haven't ran into it yet but I feel I haven't ran into it yet but I feel I'm in danger of that, cuz I've got multiple clients. Right?
And then like, you know, everybody every company's got like a partner panel, or like a support panel, or like they're named very similar things. So, if like I had two different clients, and I was like logged into uh you know, linear locally, or it pulled from Jira. And I'm like, "Oh, fix the partner panel." And there's like an ambiguity there, and it kind of like looks to what I'm connected to.
Seems like I'd have to like log off, and like log into a new uh Mac user for each session almost, to make it not pull from environment variables that might be related to a separate client, or something. Yeah. Yeah, I mean, that's what I've been using VMs for, uh which again eats a lot of resources. You need even a decent machine to do that. But, I think there's there's some middle ground, too.
I don't want to get too far into solutioning, but one thing I've seen that's effective is one, I mean, obviously you have to not have your own environment variables. That's just a problem anyway, so having a separate Mac user does does help uh when you're running those flows. But, you can set like the MCP variables like in a specific like folder, and then it only will connect to the correct like linear space.
That was the trick for me, is that it needed to always when I'm working when it's working in that folder, it needs to always connect to the correct tools, and only have access to those tools. That was the solution on the the tenancy leaking. Uh And then of course, like if you have global stuff, like, you know. So, I can talk about this, and I think this kind of gets into a more interesting one.
Uh if you're logged in on your machine just as a person, whether that I think a very common one for all of us, even non-developers, is going to be GitHub. So, like imagine the permissions that your GitHub has. Like mine has like admin and ownership-level permissions for several organizations. If my autonomous agent is acting as me, it has so much permission across organizations to cause havoc. So, like one thing I had to do was make an agent have its own GitHub profile.
For the sake of I had to be guaranteed that it couldn't accidentally use my permissions. Yeah. And so like in the environment where my agents run, my keys are not in that environment. That's That was another one of the principles. Right. Same thing for like your G Cloud. Yeah, sorry. Oh, well, cuz see this is again not a new problem, but it's Right. exacerbated, right? Like it's if you if you're root or admin you're not supposed to run as root or administrator on your local box.
Uh you know, even if it's your personal computer, you still shouldn't be root or admin level. You should have a second You should have a regular account and then an admin account. I mean, not that anybody does. But in theory, you should have that. Right. Imagine if Imagine your AI is a lower-level employee that has temper problems. And then like reflects on Am I happy that they could delete the repo if they got mad at me? for firing them. Treat it all as a security concern.
But like you said, Jeffrey, it's these are These are old problems anyway. Like the whole principles of like the the least amount of permissions needed to get your job done is like a security thing. Not even if the employee does something nefarious, but if they got compromised and someone was able to take take over. Right. In the practice.
I have not had this AI problem myself, but I have in the past had you know, a work computer and being remoted in in the good days before everybody gave you a laptop and committed code to the wrong repo because I was not where I thought I was. Yeah. Mhm. Uh and it's just like, whoops. been accidentally logged into the production database and thought you were in the dev database? I have not, but almost every I I have I yes. Almost every DBA I know has that story of dropping production. But I have not uh done that myself. Yeah.
And again, it's the same it's the same problem, right? Like if you give yourself the keys to do some dangerous things, you may find yourself in a place you didn't mean to be. Uh and I think the same is you should have even a a stricter view of Right. of of agents. And And I think it's not just about even not trusting the LLMs themselves, but I think if you have if you watch any amount of security news, you see all of these uh supply chain attacks, right? Mhm.
Like if your agent is pulling down NPM packages, and I've seen it it's very common for especially vibe coders to have uh unpinned uh semver in their package.json, so they're just pulling in like the latest package like Yeah. every time they install, which is every time they do anything on CI or do a build, they're pulling in that latest potentially compromised package, because that's when the supply chain attacks happen. The package gets, you know, released, and then people download it through these autonomous workflows.
Even if it's not an agent, like you're just running NPM install, it's going to put that compromised package Right. There's a window before it gets noticed and all that. Right. Before it's patched. So, if your agent is running against that, and now you get injected into that prompt, it's going to start exfiltrating your data off of your machine.
And if you don't have anything to catch that, or you're not in a contained environment, that could be your GitHub keys, that could be anything on your keychain, that could be whatever stored in your passwords. Like all of this data could be essentially just taken from your machine. And not to mention whatever else it could get the agent to do on your machine, especially if you're not savvy to like what's happening, you know. Oof. Yeah. You're giving me fears, new fears. Yeah, this is not funny, Dustin. These are supposed [laughter] to be fun.
I started doing VMs. Hey, I I told you. That's why I started doing VMs. [laughter] You know, you got to have some level of containment from these things, right? Because there's Imagine the same that we're all now vibe coding all of these solutions and and working at breakneck speeds, you don't think the attackers are doing the same thing? Yeah, this is kind of what you know, the idea of plugins getting up compromised is one thing, but plugins of plugins, too.
Like when you have nested dependencies and you were good about, you know, taking care of minimizing the plugins that you needed and making sure they were upgraded, but the plugins you installed had dependencies and those people didn't. And so, now you're still pulling in, you know, Right. nested dependencies that have the issue and this has happened, I forget which plugin it was, but this happened like, you know, 2 months ago or so.
Where everyone was like, "Okay, everyone we have to archive our repo, make a new one because it could have been compromised, but we're not sure. Rotate all of our environment variables." Yeah. Yeah, I've seen that across multiple repos in the last few months. [clears throat] And it's the same thing. It's It scares me, too, right? And I'm working and I'm responsible for client workloads, so I feel responsible to take care of those workloads.
But I again, I think as long as you apply these principles, you can you can protect yourself as much as you can. It may sound scary, but I don't think that it means not using these things, right? Cuz you're still You're still exposed. Like it's not like AI itself is what's exposing you to these attacks. Like you're exposed anyway. Your developers could expose you. I've been exposed to several hacks due to contractors that just didn't mean to, you know, so it can definitely happen. I mean, there's a reason that the that the thing has a name.
It's cuz it's not new. Supply chain attack. Yeah. Yeah, exactly. Yeah, exactly. So, back to our I guess back to our topic. I think tenant isolation, I think we talked we kind of covered that one really well. Um connectors, I think it's kind of the same thing. Like be wary of the connectors that you have, whether it's your Figma, your linear, you know, it's very very easy to accidentally pull in. And I know a lot of people have multiple Figmas that they have access to, not just the single client or the single work space that they're attached to.
Like it's a multi-work space account. So when you give your agent MCP access to that Figma, and to your point Isaac, you say go fix the portal, man, who knows what portal it's looking at. You know, you need some type of verification step in there or some type of isolation. That's what's drawn me to giving them their own accounts. So I can control that access. Like you really only have read only on this one work space. And then I I know for sure it's them. And you can use like the email trick, too.
So you're just using like one email plus client name or plus tool or you know, and that that makes like a systemized um isolation. That's a a good example, actually. I I talked about it in a prior episode, but I've run into that problem specifically, and it did kill production. And unfortunately, it was a side project that was like a hobby thing for me and like two people saw it go down, you [laughter] know? But it was a good thing to run into early on so I don't cause that, you know, in a much, you know, worse case.
But essentially, I had um Vercel connectors set uh on two different computers. And I had moved a I was trying to push up all the work that I'd done locally. Uh there were some stuff that was kind of like because I was tinkering around on my local machine, I had some environment variables set on my local machine and some other stuff and processes that were running on that computer and I was trying to make it just all contained so I could clone it down on a separate computer and it could still continue working.
And so I cloned it down on the separate computer and then told it to, you know, boot itself up and run migrations and set things up. And on my old flow, I had the environment variable, the local one, already populated with values and on my new machine it didn't. And so it looked to me like, oh, I could see it thinking through and he's like, I've got, you know, no local variable set. Ah, I see that this is a Vercel project. Ah, I see I'm connected. Ah, I see I can pull these values from production.
And so it actually pulled down the value for the database, ran migrations, and I had a pinned version live. So, the version I pulled down from main was not the same as the one that was live. It ran migrations, and then columns were not there that needed to be there for main, which had not been released this different version. So, production goes down cuz there's a Roll back. Then you had to roll back. [clears throat] Well, I didn't roll back. I just pushed through and fixed it cuz it was a side project. Yeah, I fell for it.
Because I mean, but yeah, I could see if this happened in the professional world Right. and I did that and brought down pride, that would be a huge um an incident report to fill out for that. Yeah, and a lot of egg on my face. So, you know, I was glad I was glad I ran into that the first week of agentic programming. I'm like, "Ooh, okay. I really need to understand the power of, you know, it will try and figure things out and even if you didn't tell it it could do some crazy stuff." Yeah, or at least then, yeah.
I've got a one that's on the environment variable one that I've seen that actually shocked me. And I I was like, "Wow, I can't believe you just did that." Like I even I didn't think you would do that. But it I don't remember what I was having it do one time, but it it for some reason it read the .env and then put the entire .env into the chat. And then like literally the next message it was like, "Oh, oops. I just exposed all of your keys. You'll need to roll all of those." And I was like, "Thanks. [laughter] That's a funny one.
I'm so glad that I got to do that right now. That's what I needed in the middle of this project was to just stop everything and roll the keys across the board, right? Like that's exactly what [laughter] I wanted to do right now. [clears throat] And it to be fair, like right, so to give you like context, like this is a platform that has like a legacy platform, a new platform that's out. They have shared environment between them. It's kind of like in a transition period. So, that's why I have those credentials locally.
Uh and it's just like, man, I can't believe you just did that. Like I that's just cost me so much work. Uh so, rolling keys due to exposure, that's uh I got to imagine vibe coders and like everybody is experiencing it being exposed to that, uh which is dangerous.
Um then the other thing I've seen that worries me a lot, this happened in G Cloud, but I I could see the exact same thing happening in Azure or AWS, which is um as the developer, you may have and may not even realize it, is what happened to me, may not realize that you're logged in to the CLI of a specific tool. It could be even Stripe that you don't realize you're logged into. Uh and the agent has access to that if you're not logged you know, if it's not in a VM, it has your your accesses.
Uh so what this one did is it went in and started looking at in you know, production infrastructure to like answer my question. Like what I was actually asking it was like go look at the docs, like read the documentation on how we built the infrastructure and remind me you know, what this how this part works. And it's like, well, I'm just going to go check the actual infrastructure like, you know, just just scanning this account and I'm like, wow, this is it's got full write access. This is crazy. Ooh. Uh so I stopped that process.
And what it made me think about is like, imagine if you were giving it a problem and it just started spinning up infrastructure for it. You could essentially have it build out, especially if you don't know what you're doing to our last episode, uh you know, you could spin up a bunch of scaling infrastructure for a platform with zero users. I actually saw something like this recently, not this bad, but a smaller case of this. Uh so you've got all of this scaling infrastructure and you have no users.
And you're like, the biggest the platform could ever be is a single box anyway. So you've got multi scale, you know, multi box load balancer scaling infrastructure in dev to test it, to make sure it works, then duplicate it in prod. So you're essentially running anywhere from 8 to 16 small boxes to do something you could do on one box in dev and one box in prod. Right? Because the AI recommended the best practice for scaling. Or or you did maybe. Maybe you said to it, "Hey, uh there's an issue where this thing keeps running out of resources.
I want you to fix it using best practices." And you're really vague. And it goes it spins up all that for you. Cuz that that was the first thing for me being logged in locally. Cuz that um if you're logged in like even AWS, that's the scary one, too, where you if you log in via the client, it saves those credentials locally. And sometimes you you'll the login, if you're smart, you've set it to expire in a short amount of short window and you'll have to keep re- re- validating every 30 minutes.
But the damage you could do in 30 minutes if it stole those credentials and you put them in an area that was already accessible to it. Um to spin up that would be very scary to me. So, again, thanks for all these new fears. Yeah. You're welcome. Well, but they're not new like uh you reminded me uh of a news article I read years and years ago where somebody had hired a company to move their car. Right? And they didn't take down their transponder.
And like they paid a flat rate of, you know, come pick up this car and move it here and deliver it by such and such a date. And this car carrier went all over the country picking up cars and dropping them off. And this guy got a $400 uh transponder bill because his car just, you know, was going here, there, everywhere. And the car carrier was like, "Yeah, you know, we we put the car on there. We charge you we charge you right We didn't say we'd go direct. We just said we'd pick it up and we drop it off by this date." And that's what they did.
But, you know, they the car it went here, there, and everywhere. And this transponder was just like racking it up cuz And it wasn't even for anything cuz it was the the car carrier was paying the toll. Mhm. Right. That's wild. [laughter] Right? Things But it's the same like it's not a new problem. It's just making existing problems worse. Like, "Hey, you know, if you your car's inside another car and it That's kind of scary. I guess it's like imagine you on your worst day, like the worst thing you've done. Mhm.
And imagine being able to being able to do that 50 times faster. Right. Yeah. That's that's scary. [laughter] Yeah. Limits. We had to put in these limits. Zero trust is the buzzword, right? Yeah. Right. All the security people are like, "Hey, hey, hey." I told [laughter] them. I told them. I told them, man. It's all It's all we got to Anyway. All right. Any more good good AI stories or should we put a pin in it?
Well, for the sake of getting this one out there cuz I said it again I in a prior episode, but it's worth saying for anybody worried now at this point about other things it can do. Um but I had a case where um I forget what Google Enterprise API I was using, but it was it was something to do with pulling like um uh a location like like a venue and data about the venue like the contact details and whatnot. Um if you pull in basic information about it, you get billed at one like cheap rate.
But if you pull in some other details like I think it was the like the image like if somebody uploaded an image or any reviews about any any reviews somebody left for the business, some of them are billed at a different rate and they're considered part of the Enterprise API for that. Uh and I discussed with uh AI which one I wanted to use and it told me what the rate would be. And I was like, "Okay, this is like $40 a pull of the data that I need." But it actually included one of those columns not realizing it was part of the Enterprise API.
And all Google does is just bill you for the API the Enterprise rate if you include that in your query. That's really all it is. Um and so I suddenly uh I had this $200 bill on Google Cloud. And fortunately it was on one of those accounts where I hadn't attached like a card yet and it's just like they they they grant you like $200 to get you used to the API. Yeah. And so it was just like exhausted immediately uh from a loop it ran. And I was like, "Oh, woah, what?" And then I asked it why it did that. And it was like, "Oh, oops, sorry.
Didn't realize that was part of the the Enterprise one." But I just imagine if somebody had their you know their debit card on file and I went to bed and woke up with a 10 grand bill, you know, that I could see that happening. So, I'm sure that happens all the time. I I accidentally had a limit, right? I accidentally had a $200 one. But, I have also seen it try like I I I got a random bill one day. It was like $200 bill for um Cloud.
And I'm like, "Huh?" And I went back and thought You know, I'd already paid that monthly I I do the 20X plan or whatever. So, I'd already paid that fee. But, when I looked it was just an auto renew. And it was like doing a dark factory pattern trying to you know, keep using keep building things in the background and I had it like auto renew. I'm like I I went in there immediately and just turned it off. And it was doing what I needed to do and I would have to pay for it anyway, so it was fine.
But, I thought, "Oh, I want to make sure I don't let it think, 'Oh, this is the best way I should bill it, but it's going to be expensive, but that's what the user wants.'" And then it just starts like I just suddenly see like 50 separate $200 auto reload charges when I come home one day or back from vacation or something. So, [laughter] another limit, another security limit we should place on ourselves. Turn off auto reload.
Well, and again, this is uh not a new thing of I have seen developers crush computers because they didn't put a limit on a SQL query and they pulled in 16 million rows. I just Well, did you know, did you need 16 million rows? Maybe you should put a limit there of A literal limit, yeah, yeah, in that case. Yeah, put a put a limit of 10,000 even if you don't think you need it. Right. Yeah. I imagine the amount of runaway queues that have been developed in the last 6 to 12 months. Oh god.
[laughter] There's got to be Last two weeks, several stories, you know. Yeah. Agnostic programming for runaway queues. Oh god. Yes. That's a bill you don't want to wake up to. Imagine that plus you also have something that hits tokens, like an API key. Mhm. On a in a and then in a queue, I mean, you could really rack up serious bills doing nothing. And it could just be a runaway infinite loop process of processing nothing over and over again. And then you add auto scaling to that, you can really rack it up.
I'm going to be telling all my clients, do not give me access. If you do give me access, give me read only. Read only. Yeah, I don't want to screw up anything for a Well, but again, it's not a new thing. Like always give the least amount of permissions necessary. Yeah. No, but but the the amount of damage you can do that fast is a brand new thing, absolutely. Yeah. Awesome. All right. Well, thank you all for listening. I'm Jeffrey Sherman. Dustin Ray. I'm Isaac Askew, and this is Never Rewrite.