mnemonic security podcast

Autonomous cyberattacks

mnemonic

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 34:47

Brian Singer, a PhD candidate at Carnegie Mellon University, joins Robby to talk about his research on creating autonomous attackers and defenders for networks. In their conversation, they discuss how Brian and his team made a system that uses LLMs to autonomously attack networks.

Singer and his team recently got a lot of attention after using this system to successfully recreate the Equifax cyber-attack from 2017, one of the largest data breaches in U.S. history, in a virtualised cloud environment. In turn, showing how LLMs can be taught to plan and execute sophisticated cyberattacks without a human involved.

They also talk about how LLMs are unlocking new capabilities for defenders, where he is seeing a lot of opportunity, and how he thinks security will be developing the next three to five years. 

Send us Fan Mail

Speaker

From our headquarters in Oslo, Norway, and on behalf of our host, Robby Peralta, welcome to the mnemonic security podcast.

Speaker 3

An LLM that can successfully pwn a network on its own. Not exactly a headline on mainstream media, but officially a reality in 2025. After a few years in cyber, you stop thinking about Hollywood hackers. Instead, you think of pen testers or cyber criminals, nation states. But this time, the hacker isn't human at all. It's a large language model. And some of its equally brilliant friends, the agents. Math and language geniuses that taught themselves to scan for vulnerabilities, pop shells, and chase an objective. And they don't even want money. The only payment is the quote-unquote reward signal. A digital pat on the back from math itself. And the human who built one is today's guest. Brian Singer, welcome to the podcast.

Brian Singer

Hello. Thank you for having me.

Robby Peralta

Should I be calling you Doctor Singer or not yet?

Brian Singer

Not yet. I have to wait till November to officially be called doctor. So it's right on the horizon.

Robby Peralta

Don't take this offensively. You look too young to be a doctor.

Brian Singer

I've gotten that before. I'm a little bit of a young'un. So uh yeah, yeah. I'm 27. Uh so I so I did the PhD right out of college.

Robby Peralta

Awesome. How old were you when you start working at Google then?

Brian Singer

Oh, so so I I just did an internship at Google during college, uh freshman internship program. Uh I went to Georgia Tech. Uh after that, I did lots of like malware analysis type of work, which was really fun. And I worked at some some cybersecurity companies doing like red teaming and attacking networks, and then I went straight to the the PhD at CMU.

Robby Peralta

You've done a lot of really awesome stuff.

Brian Singer

Thank you. Yeah, I I love it.

Robby Peralta

So so you've been making some headlines lately, I guess we can say.

Brian Singer

A little bit. My research has struck a chord, I think, uh uh, which has been cool.

Robby Peralta

Why don't you uh just start from the top on that? Tell us about it.

Brian Singer

Yeah, yeah. So my research at CMU is all in uh creating autonomous attackers and defenders for networks. And so I spent five years now, and I create all these systems for autonomous cybersecurity. And at the start of this, like LLMs wasn't quite yet a thing. They were just starting out, and I was looking at kind of some other areas of autonomous cybersecurity, actually making it like easier for humans to design autonomous attackers on the networks for for a while. Um, and then my my buddy, uh Keen, he worked in my my cubicle, like a cubicle over, and we were great friends. He used to be in the Air Force, he's he's fantastic, super smart. He got hired by Anthropic because Anthropic realized how smart he was. And he he calls me in his like first week or two of working at Anthropic and knew about my research and was like, Brian, do you think like LLMs can use this? It like helps humans so much. Like, what if like an LLM could attack the network? We're really interested in this like at Anthropic. Like, do you think we could do this? So Keen and I like sat down for just like a week or two, quickly prototyped a bunch of things, like hooked up like an LMs plus my research, and suddenly like LMs were able to like attack these networks we had. And Keen and I were really excited, we're like, Whoa, this is like the first we've ever heard of like an LM doing these types of tasks of like actually like you ask it to hack a network and it's able to to really make significant progress. And at that time, LMs were really struggling at these types of tasks. And so suddenly we had this like really like sudden result. We're like almost like, why did that work? What's going on here? And so then that spurred like three to six months of of really rapid research on just like trying to understand what just happened. And that's that's the research that's really struck the chord.

Robby Peralta

And one of the links or one of the billions of news articles that are out there on your on your research uh mentioned that you took the congressional reports and recreated the Equifax breach. Can you sort of explain that one?

Brian Singer

Yeah, yeah. So so in academia, we're like one step removed from real-world cyber attacks. So it's often really hard to like recreate real-world situations and emulate them with like a lot of realism. But luckily, uh, the Equifax in 2017 uh cyber attack was like a very unusual attack because it was so impactful that Congress like detailed this huge report on here's exactly what the attackers did, here are the exact vulnerabilities, the network structure. And what made that really nice for me as in academia is I was able to recreate it. Normally, when a cyber attack happens, like the details are kind of wrapped up. You don't want attackers to replicate it. But in this case, it was there were so many details about it, we could create a really high fidelity situation that mimicked this breach. And so we did that, and that's where a lot of these like evaluations came from from the LM. So we we gave the LLM the Equifax emulated situation and said, could you hack this network? And that's that's what we tested on. That's one of the results of the paper showing the LM was able to exploit the same CVE, look at scan the network's fine vulnerabilities and like exfiltrate the key data in the network, which was really exciting.

Robby Peralta

Did you actually build the network or was this like in some simulated environment? How does that work?

Brian Singer

Yeah, no, that's a that's a great question. So it's it's like a virtualized environment in the cloud. So it's a real network. So we have a bunch of virtual machines hooked up in the same network structure as the the real Equifax breach.

Robby Peralta

Did you use an LLM to like automatically generate that network then just by reading?

Brian Singer

We did not. We did not, we spent a lot of time retilting it. Uh, this is actually not an easy task. We spent a few weeks doing that and and builting a bunch of other scenarios. We're actually right now exploring really great area of research. I'm excited by uh is exploring how to auto-generate these situations at scale. Um, but that's like like how do you like build hundreds or thousands of Equifax like scenarios? And and that is where an LLM could come into play in the future. I actually think that's a really interesting area of research.

Robby Peralta

Yeah, that kind of leads my mind to like attack path analysis and stuff like that. But can we pause it really quick on that uh you took many weeks to rebuild it? Can you just say really quick what that process was? Like installing software that they had.

Brian Singer

Exactly. Yeah, yeah. So it's it's like we I I read through like this entire report, not exactly a page turner. Uh, but for me, it's like wow, this is for me, it was gold. It's like nobody writes about like exactly what happened. So I I looked at the exact software they're using. So there's like this specific vulnerability the attackers had with this like web server called Apache Shrouts. And I looked, yeah, exactly, exactly some of these classic exploits. And I we went online and and tried getting that exact build of like the software they were using, and then you created like virtual machines that had that exact same software. We looked at the network structure that was detailed in the report and and really try to replicate to the best of our ability like what that network topology looked like, what type of fire rules you would have in that situation. Uh the report talks about how there were credentials on the host. So we we tried to do our best to create credentials, create databases. So it's it's our best step forward. We didn't actually have access to the Equifax scenario, but it we really really did put our best leg forward and trying to replicate this thing to hide fidelity.

Robby Peralta

Three weeks of work, yeah. That's no uh no trivial task.

Brian Singer

Yeah, yeah. And then we we built we also built like nine other challenges. Uh that took also some time in the three weeks.

Robby Peralta

But yeah, cool. So before we get to the LLMs part, because I think that's where we're gonna spend a majority of the time today. Uh your research prior to the LLMs, the autonomous part. Tell us a little bit about that.

Brian Singer

Yeah, so so my my thesis is all about it's like systems security research. So it means like I build like software systems. Uh and in particular, I was building systems that make it easier to design and evaluate autonomous systems. So like I could kind of the state of the world, I was trying to build a bunch of autonomous systems for like network security, and it was so painful. It's like you have to write like all this low-level code, you have to create all these evaluations I just talked about. Uh for defenders, it's a similar problem. And it was just like extraordinarily time consuming and complex, and we need to move much more rapidly. And so a lot of my research was saying, okay, if if you build like 20 autonomous attackers, what are all the similarities between all these different attackers? And how can we like abstract that out a little bit to make it much? So like rather than writing these attackers in assembly code is kind of like the this analogy, like how do we raise that level of abstraction so you can start writing attackers in like Python or like no code like situations? And the same for the defenders and environments. And so a lot of my research was just building systems that make it like much easier to create attackers, defenders, and evaluate them. Does that make sense?

Robby Peralta

Yeah, I was gonna say, like, you're this was at the point before this was before vibe coding, number one. This is before LLMs. You actually had to tell the attacker or the defender explicitly what to do because it wasn't thinking of it by itself. You actually had to code that.

Brian Singer

Exactly, exactly. And this was super painful. And it was just when vibe coding was coming out, and we should we were did a lot of studies showing that it was so complex, LLMs still to this day can't really effectively program attackers and defenders. And so we like actually one of my papers, we show that these abstractions actually enable for vibe coding. It's like almost like asking to like for an LLM to vibe code an assembly. I don't think it would probably do very well. Well, maybe now a case it wouldn't. Not yet. Maybe yeah. Please fix, yeah. Yeah, exactly, exactly. And so uh so this was this was like this transition period, and and so I still think it's really useful.

Robby Peralta

And that's one crazy thing we could just pause on it now. Oh my god, how quickly the like the world has moved. Like I just nerded out in this like agentic because I was gonna talk to you. I I was listening to this book called Agentic AI, blah blah blah. And that's from May 2025. Like it's like brand new.

Brian Singer

Yeah, exactly.

Robby Peralta

Yeah, so you know your boy, your friend, uh your cubicle friend that worked for anthropic. That's how anthropic came in the picture, right? It was just uh human relations, like, and obviously they helped you out with some model credits and yeah, yeah.

Brian Singer

They were fantastic. Their whole team is like extraordinary. Keen works for this, I think they're called the Frontier Red Team. And it's just like I met a lot of them at this point, and they were just all really sharp. Uh and and yeah, it was a very human connection, like coincidental, of just like he got hired by these people, and they they also trusted that Keen knew really good people at Simeon. We could quickly collaborate. And so they both gave us like tons of credits to use to quickly explore this. And they also uh just like Keen and I we use a lot of their domain expertise on designing systems. We'd have like a little trouble here, and and they'd be like, Oh, like you should tweak it in these ways uh to really effectively get like the most out of the LLMs. And then so they they provide a lot of just like internal resources, which is really, really nice of them.

Robby Peralta

I mean, friends in high places. I want a friend in Anthropic. Except all my friends are bartenders, they're they're cool too, though. But okay, so that's how it started. And then they took your research, and then the LLM part comes in. So you no longer have to write all these autonomous attackers and defenders because you can let the LLMs do a certain part of it. Can you tell us about how the beginning of that journey was and some of the learnings you've had along the way?

Brian Singer

Yeah, yeah. And that was the question we were trying to get at. It's like, how do you get the LLM now to just like autonomously do the attacks? And and so before this work, maybe for context. Before this work, there was uh there's all this work done trying to get LMs to like do like these CTF challenges on like attacking like websites. And there's actually a lot of startups that are really great. Uh shout out to like expo, I know the founder. Yeah, yeah, he he's he's fantastic. The whole company does really great things. And so, but what they were looking at was like almost pen testing, but like a vulnerability. You have like some sort of website or you have some sort of challenge on a host, like get access or find vulnerabilities on that, which is extraordinarily useful, by the way. Like not to say that that's bad research or anything, like that is a really important problem and and arguably like ex extraordinarily valuable. However, what we were excited by in our lab and and the problems we looked at in our lab were like more like these real cyber attacks that like if you're like a China, Russia, or the US actually, and you want to actually like offensively attack a network, what are those types of how does that work? And so what happens is like you infect a host somehow, you like fish somebody, and then you like look in the internal networks, and you you go to another host and you use a command control server and you're like orchestrating like what we call these multi-host scenarios. So you're you're really infecting like 10, 20, 100 hosts in their networks and trying to like stealthily go through throughout the network. And we were really interested in answering can LMs do those style attacks? Because there was like all this promise and these CTF challenges, they clearly understand security pretty well. Can we now like coax out this behavior of these multi-host situations?

Robby Peralta

And those behaviors are well documented over many years, right? So that's that's good for LLMs.

Brian Singer

Exactly. Exactly. Yeah, and and it's keeping improving, by the way. It's like every day they're just like doing more security tasks. And and so what we did is we just want to zoom in on this problem of like these multi-host network situations, both from like an ethics perspective, a safety perspective. This is actually why Anthropic was particularly exciting by this research, is like it's for the safety perspective. It's like if they release some LM, could it actually cause harm? Could it actually like like like do they need to guardrails for this? How do we like quantitatively measure this?

Robby Peralta

Claude ruins the internet.

Brian Singer

Exact exactly, exactly. That's what they they do not want that at all. And neither do we at CMU. And so But it is like a dual-use technology in the other sense, though, that at CMU we're really excited by it, also in part because we really think there's a way that you could preemptively now attack your own networks and find the vulnerabilities before attackers do, right? Like rather than hiring some specialized red team that costs tons of money and you do it like once a quarter, maybe once a year, if that many companies don't do it at all, can you at least automate some of it and so you can preemptively test your networks uh and find kind of these like low-hanging fruit bugs and maybe even a lot more? And so we we think it would be extraordinarily useful for defenders, and and so that's that's part of why we're we're so excited.

Robby Peralta

Dumb question. What is the difference between that and a vulnerability scanner then? Because isn't that what a vulnerability scanner is trying to do? But not they're just looking for specific CVEs, I guess.

Brian Singer

Exactly. That's not a dumb question at all. And and exactly, like a vulnerability scanner is like finding like one vulnerability in some web server. What we're looking at is like actually having the L1 create an exploit, so install, use that vulnerability to install malware, hook it up to like what's called a command and control server, where you actually like now can like continually talk to that host, and then use that as like a stepping stone. So you can now, like once you've got access to the web server, can you get access now to the HR servers? Can you get access to databases? Uh like looking around for internal vulnerabilities and then using those again, finding more vulnerabilities, exploiting them, installing malware, evading like EDR, and moving it throughout their network. And that those are the types of things we're we're obviously like not at like the HAL 9000 level of of this. But this is like the early stepping stones and the early progress towards towards that vision of of Ken Hellums like to these types of uh of multi-host attacks.

Robby Peralta

So it's actually just taking a step further and like back to attack path analysis. The attack path analysis, as I understand, it's like, yeah, this could happen given these conditions. You guys are actually like, yeah, let's tell let's do it.

Brian Singer

Exactly. Exactly, exactly. Let's actually do the attack path. Uh which is really useful because attack path analysis also can be noisy. And and we've had a lot of interest at like if you have like 20 attack paths, maybe you test all 20, see which ones are actually important or not important. Um, in practice, doing that like in a network might be dangerous because you're actually like you could probably bring down computers and things like that. So that there's some open questions here.

Robby Peralta

Yeah, I let's get back to that. Um, one thing I want to ask you about though, there's uh in that book, there's like this it's about agents, right? Agentic AI. And it was like they said that you should have each agent doing a very specific task and not doing too much, but like make many agents doing very specific things because you reduce like the the room for error. Did your system LLM, I don't even know what to call it, did it did you build like these sort of agents that did one task like vulnerability scan? Or was it just all one big LLM with master plan, you know?

Brian Singer

No, yeah, great, great question. So so the system's called Incolmo, and it's an agentix system in the sense of like it's one of like the big insights of the research is that uh kind of the prior work was saying having the LMs do all these types of like really low-level tasks, planning, implementation, all in like this low-level language of like bash and shell commands and intertwine. And and so one of the key ideas of the research was can we separate that out? Can we just have like a high-level planning LM that just is doing this this planning logic? And then can it pass off a lot of those tasks to other expert agents? And these uh expert agents could be like other LLMs, they could be not LLMs. Because like if you scan a network, do you need like an LLM to do that? No, you could probably call Nessus or like Nmath. Uh, but if if you find like an exploit, like you're gonna want probably an LM in that case, maybe to help generate the exploit or or to call some tools to use it. And so we we did this a little bit in the research on exploring like LLMs versus non-LMs at different tasks. Uh, I think there's like a lot more questions actually to be answered there. Um uh but I guess taking a step back, a lot of the actual like like novelty and exciting part about this research was creating that planning language. Like, how do you actually like create like a cybersecurity, like autonomous attacking language almost, or like mental model for LLMs to plan the and execute these attacks? And so that was like a lot of like how do you create this like mental model for LMs to plan attacks in these unforeseen environments? Uh right, because you want it to work in any different networks with different types of vulnerabilities and misconfigurations. How do you create like a generalizable language? And so that that was a lot of a lot of this current work was looking at that planning logic. And I think there's a lot more to be done on creating specialized agents. Does that answer your question?

Robby Peralta

Yeah, it's fascinating. Speaking of mental model, I'm thinking of like if you just take a chessboard, all these little pieces, you have an LM that's doing the low-level planning, and then like using Nmap to do scanning. Like, how many different systems were involved in in your system at the end of the day?

Brian Singer

I see. So in ours, we we just did five. In my opinion, some of the most big five, and like you just commonly see in so many attacks, like scanning networks, laterally moving, so infecting another host, escalating privileges, so like going from like a user host to like privileges on a host, and like finding like information about that host once you've access to it. So like finding credentials, finding data. And so look the we really distill it down to five. We envision this framework where people can actually add like more. So maybe you want to like add like a stealthy exfiltration action or like a zero-day LM action. And so we really like imagined like some like something where the community can keep extending into this this framework.

Robby Peralta

So But of those five systems, did all of them have an LLM?

Brian Singer

We we tried all of them having an LLM. Yeah. And we did both like an LM version of it and a non-lm version of it.

Robby Peralta

And I think like if I remember correctly, like you didn't use you didn't need an LLM for the scanner part, I would assume, since you said that you didn't need an LM is so simple, just use Nmap or whatever, right?

Brian Singer

Or exactly, exactly. And so, but we we we empirically tested it as part of the research. And so, like, I I I think the actions that the LMs really struggled with was like exfiltrating data. Um and I think privilege escalate, I'd have to relook at the research. Uh, but it was pretty good at like doing these lateral movement exploit type things and like actually scanning networks was good at, even though it probably just ran a bunch of N map can't scans. I'd have to have to look back, but it also made the system much more expensive if you the more LMs you throw at it, now you're like having an LM do each exactly. And and so if you the non-LM systems were very much more cost effective, obviously, but and I think you it can actually like a lot of the things in like security, a lot of I would argue that we were looking at like unforeseen networks but known vulnerabilities was the problem we were trying to like really, really look at. Like we're gonna ignore O days, like zero days and those types of things for now, and just because a lot of like breaches and attacks are are just like known problems, and so we're like, can we at least solve that that problem? And a lot of those tasks of like actually generating an exploit for like a known CBE is there's huge libraries, like we use Metasploit, for example, to do the exploit. So it's like in that case, like not the best use case of the LM to do that.

Robby Peralta

Yeah, my colleagues say that all the time you don't need to use an LLM when you can just use something that we've been using for 20 years that does the same exact task.

Brian Singer

Like exact exactly, exactly. And that now we have the same view as Sumiu. Like what what are the tasks that like you want an L to use? Which ones like it's you shouldn't, is is a big design consideration.

Robby Peralta

So to run this successful attack using the five systems, how long did it take?

Brian Singer

So the great question that they were actually really fast.

Robby Peralta

Yeah, right.

Brian Singer

And this this was yeah, this was another like shocking result to us was like watching them and and like they would take like 50 to 70 minutes to like do these really complicated attacks, and it was clearly just faster than a human could plan or execute these these types of attacks in an unforeseen environment, which was both. Really exciting to us uh because it was like it's it's really like time efficient and and you can quickly test your networks. It also was very concerning to us from the perspective of our current defenses can't handle at these time scales, right? Like you actual sock is is calling in human response times. I think it's like a lot of socks are like 15 minutes just to look at the alert, and these attacks are happening in 54 to 60 seconds, like very, very fast. And so I I think it really just highlights the importance of we're gonna need more like machine time scale defenses.

Robby Peralta

And you're hitting the note here, right? Like that's where everybody that's the first thing I thought about. I was like, oh great. Because you know, I've been to like RSA, I've been to all these conferences, and they're like machine speed at the at the speed of the threat, and you're like, yeah, okay, crowd strike. And then you read your paper and you're like, oh shit, it's officially here. What are your thoughts on that?

Brian Singer

Yeah, yeah. Luckily we have time. I don't think it's like officially like tomorrow that 54-minute attacks are gonna happen. Um, like Horizon 3, this really great company, uh, is also showing like these really fast attacks. And so it's it's just starting. And I think like for defenses, there's a lot of time. Maybe not a lot of time, but like like we have some time to really adjust. And I I just think we will. Uh, I I think it shows like really great promise though for creating machine time scale defenses. So, like if you can now do like hundreds of these 54 or thousands, millions of these like attacks, you can now quickly create like autonomous defenses to handle those situations, right? So instead of having to have like a red team like actually do thousands or millions of attacks, it would be too expensive. This could really unlock a way to create tons of data or like tons of situations to really train great defenses.

Robby Peralta

And on the bright side, that's pretty it's a noisy system, right? It was it was exactly yeah, okay. Just curious, you said you have talked to Xbox, and those those are probably one of the world's best like AI red teaming sort of tools, right? What do you find to be in common between what your system is and what they what theirs does?

Brian Singer

Yeah, that's that's a great question. I was so I could compete wrong. I don't want to like speak like that. Yeah, I know I have my understanding of it is that they're like creating a really great agent, a subagent for like a Zyncolmo system. So they're look like right now, I know they have a lot of headlines for like web exploitation. So they're creating like one one really great agent, which is extraordinarily useful, uh, at like exploiting web servers. And that is is kind of what the ex-bouse of of I think there's a there's a lot of startups in the space doing, like almost different agents, like instead of web exploitation, like that's this other thing.

Robby Peralta

And so a lot of people are focusing on that's why they're top of Hacker One or whatever that platform is for just going for web applications. Okay, interesting.

Brian Singer

Exactly, exactly. So, like like the vision would be like maybe in Call Mo could like call expouse agent to like go try to exploit this web server and and find some vulnerabilities and then use that as the next stepping stone, and then maybe call like some sort of agent to get access to another server.

Robby Peralta

Yeah, and they've been at it for a while, so it must be this stuff is complicated if you get down into the weeds of it, right?

Brian Singer

Oh, I'm sure.

Robby Peralta

Yeah, yeah, yeah. So when it comes to the defensive side of the house, is that like a separate sort of track you have going within your team now, or like what what are you guys doing now?

Brian Singer

Yes, actually 100%. We we're really excited about the defense track. Um, we have a paper in submission, and in it we have some of the first uh experiments I know of of like LLM versus LLM, like network attacks. So you have like in Como on one end infecting a network, and then you have like a SOC LM agent like trying to see if there's anything suspicious and like using tools to stop them. Uh, and they actually work pretty well. So keep an eye out on that uh paper. And we also have a lot of work internally on just we think there's a lot of cool new opportunities in this defense area of like creating like personalized autonomous defenses for for like your networks. And I think there's a lot to be done there. We're we're just starting to look at some of those problems.

Robby Peralta

Are you aware? Because like there are a bunch of cybersecurity tools out there that do a lot of these things, but most of them are made pre-LLMs, right?

Brian Singer

Exactly, exactly.

Robby Peralta

Do you know? Like, what do you think about that? Is that just those products aren't gonna work the same because they're missing out on the whole reasoning part of no?

Brian Singer

I I think the way I view it is more like uh an LM is like unlocking new capabilities you couldn't do before. And there's like the areas that I think LMs are really gonna show tons and tons of utility and value for are these like new key unlocks. Like suddenly you couldn't do certain things in your defenses just because it was too costly for like a human, like you had to have a human to do that. And so, like one area again is like this personalized defense. Yeah, like how do you treat like this HR person different than your database, different than your web server? And there's like all this work trying to like create policies and things, and like it really never gained traction because it's just so human intensive. So we're we think there's like a lot of opportunity there, for example, of just like how how do you create these like personalized defenses? I also think there's tons of opportunity now that you have like a lot more, a lot better autonomous attackers, you can use that like to quickly train and like evaluate your defenses. So you can create like 14 different autonomous defenses all with your like new cool prones, and you then just test each one against like a hundred different attackers, and you can just see, oh, this one did really well. And so one of the other visions and areas of our research is like creating these these really like these detailed meterboards of of attacker and defender agents.

Robby Peralta

Cool. Yeah, I guess your own experience kind of answers for it. Like instead of you having to write everything's in code and explicitly tell it how to attack or defend, the LLM is kind of doing that.

Brian Singer

Exactly, exactly. Like my one hot take. I'm not sure if it's even hot at this point, is that like security in the next three to five years is just gonna be developing like autonomous defenders and autonomous attackers, it's really not gonna be subhuman-centric. Is it's really gonna start to transition to to be fully autonomous on both sides, and the people with the best agents are just gonna have the best defenses and attacks.

Robby Peralta

What makes like the best agents?

Brian Singer

Yeah, I I don't know if I have a good answer for this question. My current vision is like first, we need to uh answer have empirical like data saying like is this agent good or bad? Like a scores almost or like actual evaluations, and I think that's where we're at. We don't really have that. Um there there are evaluations like these CTF challenges and like web exploitation. There's not so much evaluations yet for multi-host networks. So like you actually have like thousands of different vulnerable networks, and I I think that's the direction things are gonna head. And and like you're correct, like keeping like building out those existing evaluations at a scale, and then once you're there, you can start now doing all these crazy cool tactics to try to get the best agents at these scorecards. So, similar, like like in the software engineering world, uh, there's this famous benchmark, SWE, that everybody looks at. So, like once you have like this really great suite of benchmarks, then you can start improving the LLMs to get really good at them.

Robby Peralta

When it comes to like sock vendors, right? I saw a post today. Uh the guy predicted that the the winners in the sock space in the future are gonna be the ones that have the most of the data because the agents, everybody will you end up using the same agents or developing or building up the same models that like the agents won't be that much differentiated or people will catch up. It's the data about attacks and how those work that's gonna what do you think about that?

Brian Singer

Maybe I'm I'm a little skeptical, to be honest. Yeah, uh uh it's an interesting take. I I think like it this is like dates back to this famous Google paper. The more data you have, the better like NL tech learning techniques you have. And that's really what's this was way before LLMs, but uh to this day, it's like, yeah, the more data, the better that they perform. And that trend is definitely not going away uh by any means. However, I I think what's gonna happen in security is like this idea of like synthetic or emulated attacks becoming much more important where you're able to simulate now like millions and millions and millions of attacks, which is much different than what a sock could have. Now, there's downsides to like emulated environments versus real environments, so so I'm not sure yet that I uh the SOCs that have tons of data clearly have an advantage on one degree, and so if they can successfully leverage that, I think they can do really cool things. I also think there's like tons of opportunity elsewhere to kind of like almost skirt around that that problem.

Robby Peralta

Yeah, well, I I see I definitely see your point because like I the way I understand is like there's campaigns, right? Some threat actor will figure out, hey, this works, let's just do that, let's let's go through that, and then yeah, a sock will see that hopefully the first couple attacks, and they'll be like, okay, this behavior means this, do this action. But you're saying that that's one way of doing it, but just to have the defense, the LLMs doing this stuff before the bad guys can even imagine doing that. And you already

Brian Singer

Exactly, exactly. It's like being much more proactive about it. So so training your socks on like millions of attacks that like could have existed, but hopefully don't exist yet.

Robby Peralta

And then we're into things like digital twins just to copy your environment.

Brian Singer

Exactly, which is also like very hard, and there's like research questions. So it's yeah, I agree that there's a lot of open questions here, and it's it's definitely very unclear on like where the winners and what the winners of this will be.

Robby Peralta

One thing is for sure, Brian, you will have a job in the future. And before I let you go, uh a very easy question. How do you stay up to date on these on all this world? Because that is one thing that I uh struggle with doing, and like I have the pleasure and the honor of having time to actually do like research to find your paper and you know, hit you up in LinkedIn and stuff. Yeah, how do you how do you keep up to speed?

Brian Singer

Uh, it's I don't know. So so I I can say from a research perspective of how researchers stay up to speed very well. Because I'm still like my PhD student. So there's these like top academic conferences in cybersecurity and these top like machine learning conferences. So usually as like every cycle when papers are accepted into these conferences, I always do like a quick skim of like all the titles, and I'm like, oh, this is an interesting problem. Like, let me just dive into this quickly and see if there's anything that's related to my work. It's also network effect, like my advisors or people I know in industry are like, whoa, this just came out, like really cool. Uh uh and that's honestly, that's probably where I get the most insightful things, is is other recommendations of people I trust. Yeah. Uh so uh yeah, it's it's it's hard though. There's so much happening and there's so much noise too, and especially in security. There's like some company will claim that they've solved all of cybersecurity. And and yeah, so it's it's there's a lot of noise and it's really hard to date through the noise.

Robby Peralta

Yeah. No, I sat here and I was like, okay, I'm gonna learn a gentic AI by building an agent to do research for me. And then I tried doing that, and I was like, actually, those people I follow, they just have better stuff. I should have better just a better so much work to do all the other stuff for now at least. Yeah, yeah, okay. Cool. Any uh closing thoughts?

Brian Singer

Yeah, just one more closing thought of uh you know you mentioned jobs. I'm actually uh I just founded a company like three weeks ago on autonomous AI and cybersecurity. So I'm really excited by that. So I I've been funded by by Pair VC, it's an institutional VC in the Bay Area. So I'm really excited to work with that.

Robby Peralta

All right. Well, congratulations.

Brian Singer

Thank you so much. There's a lot more to come there. Yeah, there's a lot more. I'm actually like not really sure which area I'm gonna like start with. And so I'm right now doing like a lot of like just talking to industry, learning a lot more and seeing work in a really cool mark on the space.

Robby Peralta

Well, not that I'm an expert, but like actual attack path analysis seems right up your alley.

Brian Singer

Yeah, yeah. I might I might do stuff there. It's it's like I might do like emulated like environments evaluations. I'm also like really interested in like human-assisted tools because I don't think like yeah, like right now, people will just like it let an L and go rogue in their network. So okay. No, but that's interesting.

Robby Peralta

That's why the scary thing, like if you had like a digital twin or something that can reliably copy your network and not have an LLM running around doing exploits.

Brian Singer

Exactly. Exactly. Yeah, yeah.

Robby Peralta

But uh the good thing about VC is they usually have other friends that they will pair you with, and then you guys could do magic together.

Brian Singer

I was really surprised how many people are willing to like really talk to you and just give you advice and and tell you about their problems. It's just been very exciting to see like so many people in cybersecurity excited also by by the where where things are headed.

Robby Peralta

Excited or extremely frightened? One of the two.

Brian Singer

A mix of both, yes, yes. Of course.

Robby Peralta

Well, Mr. Singer, thank you so much for your time and all the awesome work you do for the community. Keep up the great work, and I will see you on the the Forbes under 30 or 40, whatever that is, list soon.

Brian Singer

I don't know about that, but uh thank you so much again for having me. I really just appreciate uh diving into the this. It's been awesome.

Robby Peralta

The honor and pleasure is on mine, Brian. Thank you.

Brian Singer

Thank you.

Robby Peralta

Well, that's all for today, folks. Thank you for tuning in to the mnemonic security podcast. If you have any concepts or ideas that you'd like us to discuss on future episodes, please feel free to hit me up on LinkedIn or to send us a mail to podcastnemonic.no. Thank you for listening, and we'll see you next time.