mnemonic security podcast

ML Engineers these days

mnemonic

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 26:46

Have you ever worked alongside a machine learning engineer? Or wondered how their world will overlap with ours in the "AI" era?

In this episode of the podcast, Robby is joined by seasoned expert Kyle Gallatin from Handshake to enlighten us on his perspective on how collaboration between security professionals and ML practitioners should look in the future. They discuss the typical workflow of an ML engineer, the risks associated with open-source models and machine learning experimentation, and the potential role of "security champions" within ML teams. Kyle provides insight into what has worked best for him and his teams over the years, and provides practical advice for companies aiming to enhance their AI security practices.

Looking back at our experience with "DevSecOps" - what can we learn from and improve for the next iteration of development in the AI era?

Send us Fan Mail

Security in AI and Machine Learning

Speaker

From our headquarters in Oslo, Norway, and on behalf of our host, Robby Peralta, welcome to the mnemonic security podcast.

Robby Peralta

Danes, Swedes, and Norwegians, developers, machine learning engineers, and security people. Same, same, but different. According to Gemini, they all spend their time staring at screens, wondering where their lives went wrong, they have a love-hate relationship with coffee, their fuel, and their curse, and they all secretly believe their job is the most important one. But we'll never admit it out loud. On a more serious note, we seem to be at the peak of inflated expectations regarding AI. Which means that pretty soon, there will probably be speak of another AI winter as we enter the trough of disillusionment. But no matter what is said on LinkedIn, there's no doubt that security has inherited a new risk vector from LLMs and AI stuff. And therefore, I thought we'd have a chat with someone that works with them to pick his brain about what security means for him moving forward. Kyle Gallatin, welcome to the podcast.

Kyle Gallatin

What's up? It's good to be here. If you're wondering uh why I'm dressed like this, it's because we have a Hawaiian party afterwards. A hula hula hula thing. I guess it's not the first time that uh you know that I'm a party guy. We uh we uh we had some good times together in Switzerland. We did, yeah.

Robby Peralta

So thank you for that. That was awesome. And um you have substance, Kyle, and not just substance, you have like AI machine learning substance. So extra points there.

Kyle Gallatin

Its my whole thing.

Robby Peralta

Yes. So you are a uh you're a published author of the uh let me write it, the machine learning with Python cookbook, practical solutions from pre-processing to deep learning. And you wrote that a while ago. You didn't jump on a bagwagon.

Kyle Gallatin

Yeah, no, I was and uh I'll say that I am a co-author. So I I wrote the second edition. First author did a lot of work, so I didn't have to. But yeah, I uh had that out last summer, actually. Yeah.

Robby Peralta

Cool, cool. Has your life changed at all since since ChatGPT came out, or is this like nothing new for you?

Kyle Gallatin

Um As someone who like builds solutions with AI, there's a lot more generative AI and large language model solutions that I'm like building into kind of like products on an everyday basis. Um, like a lot more things that I like have education initiatives around because people want to learn about these things, and that's what I do as well. Uh, but there's been a lot to do. Yeah.

Robby Peralta

Yeah, I guess one positive benefit is people understand maybe a little more about what you're doing than before. Like now they have a little bit of context, I guess, right?

Kyle Gallatin

Exactly. Everyone's involved.

Robby Peralta

You have worked with a bunch of awesome companies, uh, Pfizer, where you were a data scientist and machine learning engineer. Then you went on to Etsy, where you did software engineering and machine learning. And now you're with Handshake, uh, which at face value seems to be um a platform to help students connect with potential employers. But under the hood, I guess it just runs on a bunch of code and algorithms that you're a part of.

Kyle Gallatin

Exactly. That's a perfect description.

Robby Peralta

Is there anything else you want to add to your background?

Kyle Gallatin

I used to be a biologist um for a little bit there. Yeah. So I have a a useless master's in molecular biology. Um but yeah, you pretty much summed it up. It's all data scientist, machine learning engineer, software engineer, and everything in between.

Robby Peralta

Yeah. So conversation today is I have a a theory uh that I've been like in my head recently that like uh security is uh I think it's an under, it's not like thought about too much in like the AI ML world yet. And I want you to tell me if that's true or not throughout this conversation and to figure out like uh where we are in the states of your world and security.

Kyle Gallatin

For sure. I like it's probably not honestly thought about as much as it is in in some other places. I mean the like the attack vectors, I feel like for ML are are new in first in some ways and very different. So on one hand, it's software, so you have the same kind of security concerns as you would for building any kind of application where you want to, you know, protect endpoints, you want to expose things publicly, all that kind of stuff. But at the same time, like the ways that people are attacking ML um are kind of new, and it's not something that people often think about when they're building an ML application, I think. It's definitely not top of mind.

Robby Peralta

What has security meant for you?

Kyle Gallatin

It's meant a lot of different things um over over time because there's you've all of the classic aspects of data security, software security, and then you have new things of ML security. So for me, um, it's meant, of course, data security, which is you're working with um data. Sometimes, you know, it's of a sensitive nature if it involves users or you know, advisor their healthcare records, things like that. Um, there's a lot of like rules and regulations and specific ways that you have to access that data so that you're not putting individuals at risk or um potentially even identifying them in your own work. Uh you can't like join two data sources that might give you too much information. Like there's a lot of rules around that. Um the data needs to be accessed in a very, very safe and access-controlled way. Uh for software, it's been um classic software things. You know, it's like you're you're building an application um for a company. You want to be careful about who can access it, how they access it. So that means you know, don't accidentally build things and expose them to the public internet if they're sensitive. Don't um make sure that you have role-based access control. Um, make sure that yeah, things live within certain VPCs if they need to stay within that VPC. Um, that's what that's meant there most of the time. And then on the machine learning side, there are things to think about, but I don't think that we've thought that much about them yet, or had to think that much about them. You know, there are attack vectors for machine learning models, mostly it's around stealing IP. Like if you had access to a machine learning model or a machine learning model endpoint or something, you could like reverse engineer a model or like like get it to leak data, or there are common attacks that people do with like chat GPT to get it to like do stuff it shouldn't do. But yeah, there are some interesting attack vectors from machine learning too.

Robby Peralta

Yeah. So that before we got into machine learning, you said everything that that sounds like you've actually worked properly with security, like you named like privacy and like identity, and like just like the whole the guardrails of accessing things. Uh but yeah, just the way in your answers, it seems like the machine learning is kind of like new and you haven't had you haven't been bombarded with people, security people trying to make you do things differently yet, I guess.

Kyle Gallatin

Yeah, I think that like I mean they want us to do the same, like the do it the safe way for data and software, like in is is my experience. And that covers like 95% of what you have to think about, maybe if I'm like maybe for an ML application. But there's new things with ML that um that do open up, like new modern attack vectors, um, that are I think are going to be interesting to see all over the next years. If we saw a lot of people, you know, breaking GPT, uh chat GPT, um, I think that'll it'll be similar things for a lot of different uh a lot of different large models that get deployed over the next few years.

Robby Peralta

Your relationship with security people over the years, uh is it do you share the same sort of vibe that I gave off? Like the security people that just obviously been on the outside and just trying to like come in and boss people around, or like how is how your interactions with them been?

Kyle Gallatin

So it actually has varied. Um, it's not like like sometimes if security folks aren't like I think there just needs to be a good partnership. Because yeah, there have been times where you know, if you bring in security at the end of a product, you're about to deploy an application, they're like, dude, what are you doing? You didn't do this, you didn't do this, you didn't do this. Like, there's all these rules we have for this application, like to deploy an application, you haven't done any of them, and now the product's like delayed two weeks. And it's like, uh, yeah, I know, like super pain, but we also didn't tell you whatever. Um, at the same time, I've had security teams who are basically partners. They're like, look, like we're not here to slow you down, we're here to do things in the safest way possible, the safest and most efficient way possible, and the best balance of those two things. So I've also had security folks. I think some of the security folks I work with today are a great example of that. You know, if they're not slowing you down, they're not like unnecessary burden or risk. They're really just there as a resource to like help you build whatever you're gonna build and make sure that it's built safely without undue burden on the developers.

Robby Peralta

That's probably like the nicest way I've ever heard uh somebody that has worked with development ever portray a security person. But I think that's because like the the places you've worked, those are like those are developer-centric environments. So that makes sense that they didn't uh yeah.

Kyle Gallatin

Yeah, they're engineering teams. So like, you know, there's there is some like level of like respect for that vibe, and we'd like and you know, we're all in this together, making decisions, that kind of thing. Um is it different in places where it's you know, like like non-engineering? Like, I mean, at a larger company like Pfizer, like we weren't necessarily an engineering team. And so like I feel like there was some stress with security sometimes just because

Model Registry and Security in ML

Kyle Gallatin

we didn't work with them closely, we didn't have a pattern of working with them. Um so when they would come in, it was like uh like the here are things we have to do now. Um things we should do.

Robby Peralta

And do you feel like those uh security people, like in your recent interactions, they understand like how your world is, uh, or or is it kind of like you have to teach him something new every time?

Kyle Gallatin

I think they um most of the security folks that I've worked with in the past few years, like are pretty are pretty tuned into the like development environment and kind of get what's going on. They understand um they understand the applications we've we're building and honestly they even understand ML. So and like most of the time the folks in the engineering orgs are also well versed in cloud infrastructure too. And so it's um it's often a pretty seamless conversation where they understand honestly more than I thought they would about like what I'm trying to do and achieve and deploy.

Robby Peralta

Yeah, right. So um there's most there's pretty much only security people listening to this uh podcast, right? So if I asked you like what does ML and AI mean these days, like I've understood that there's like these like app stores, but they're not apps, they're like model stores, I guess. And then you go out and you pick and take from certain things, and you bit you that helps you build whatever you're trying to build. Can you explain that process of like the the app store for models?

Kyle Gallatin

Yeah, for sure. Um there are, I guess you'd like could call it like a model registry or maybe model catalog. I don't know. There's probably a bunch of different ways, but companies and websites, for instance, like Hugging Face is a big player in that space, where you can go onto their website, sign in, and they have folks all over the world kind of registering like large language models or computer vision models that are like pre-trained on some task and that you can either download and fine-tune for your own use case or download and just kind of use off the shelf. So, an example would be like Facebook's Llama large language models, which are like GPT-like models. You can use them for conversational things, you can use them to generate texts for a lot of different things. Um, you could go on talking face, download them, and just start to run them yourself if you wanted to like have your own local chat GPT or start training it to do other stuff for your specific use case.

Robby Peralta

Cool. And that's kind of like just out of the goodness of the whoever put those models there's hearts, right?

Kyle Gallatin

Pretty honestly, pretty much. I'm sure like yeah, for the most part, it's like in the spirit of open source. Like people just giving back to the community so that everyone can benefit from it.

Robby Peralta

Yeah. And what risks do, if any, do you see with that sort of economy and way of working?

Kyle Gallatin

It's probably pretty similar to the um pretty simple the risks you see with like any open source, any open source thing, right? You have a risk of like either accidental or deliberate vulnerabilities being built into kind of like large open source like tooling. Like, you know, tons of people are going on and downloading like llama models from Facebook. Maybe that's like a legitimate source, it's Facebook, you don't, you know, you trust it. But maybe like individuals are uploading models for more specific use cases. And you know, they could could hypothetically build little like weird little backdoors into those artifacts, or you're you're just downloading like a sometimes just a large binary file. So you could theoretically anything could happen. So it's it's kind of like the the same thing you go through when you're looking at like an open source project. Like, do I want to use this? You're like, is it safe to just pull down this code and start running it locally and just hope it does what someone said it was going to do?

Robby Peralta

But so you're aware of that problem, right? Uh you said what HubSpot, do they actually do like a S-bomb of like what's those components? How deep did like what sort of level of security do they put on for you?

Kyle Gallatin

So it's called Hugging Face, and I'm actually not not 100% sure um what the kind of like security setup is. That'd be interesting to go on there on the website and check it out. Um I'll be honest, sometimes I'm just blindly downloading large bin files that contain machine learning model weights, and it's a process I have to like load them into like a model like and do some stuff in Python to get it running. So maybe there's not as much risk as as as I'm thinking here. But you know, if you're just downloading stuff from the internet, of course, there's some risk associated with that. And I assume they have some process to be like, all right, you can't just upload this clearly uh heavy thing.

Robby Peralta

No judgment from my side, just so you know. Like, I we uh the whole point, like again, is just like to get people, security people to understand like how you work, and because we all understand, like you have deadlines, you have shit to do. We wouldn't have a job if you guys weren't out there making a product and shipping the product, right? So uh it's uh it's mostly just to understand and get everybody on the same, just uh avoid this whole developer and security clash again horse.

Kyle Gallatin

And I'm not here to lie about it either. I mean, I think that like like for many folks, like ML practitioners, especially, I'm sure security is like a little far back on like the list of what's going on. Um they're you know, the machine learning world moves so quickly. So you're constantly just like installing new libraries from the internet, you're constantly like looking at new open source projects and pulling them into your workflows, downloading new models and then trying to train those. So, like the amount of new code and files and things that are just like coming from different sources into like a machine learning practitioner's workflow, it's like pretty high volume compared to someone who's like maybe just maintaining, I don't know, some older app or something like that. There's a lot, yeah, a lot of stuff there. And there's just so many tools like for deploying Python notebooks and stuff. Like compute as a service is a big thing within machine learning world to train these big models. So you have companies that um basically will offer platforms to like easily write code and easily write machine learning, and then you deploy them in the cloud. If you don't secure those, people are gonna find those endpoints. And if you don't change the password, they're gonna log in and they're gonna start mining Bitcoin because you just deployed something that like a like a Python notebook that's unprotected to the internet. People can see that, find that, and then all of a sudden you get a notice from Google. Um, which it's great. I have seen this happen where Google emails you and says, Hey, like you deployed something and someone's mining Bitcoin.

Robby Peralta

By the way, and uh can you just give us some context for how like how many projects or how many things you download, just so we have like a pure understanding of like the threat surface of a typical day or typical week at work for for someone in your position?

Kyle Gallatin

Yeah, it's definitely gonna vary week by week, but you know, I think on the library side, you're always looking at there's a bunch of new large language model and generative AI Python libraries right now. So it'd be like five new Python libraries, just pip installing things and being like, oh, what is it like? Let me try this out. Let's see if this works. But then on the model side, if we're saying I'm like doing something where I'm comparing, like I want to build like an internally hosted large language model for a different task, might go on Hugging Face and download like five different Llama versions or five different versions of some other model and then evaluate them and see which one works the best for my task. So you like I'd say like that would be the threat surface, and I'm not sure if like five libraries, five models, that's uh quantifiable enough.

Robby Peralta

Yeah, but a lot of those like uh libraries or models, whatever, they also have like underlying things that go to other places, right? So like there's a tax surface which is each of those.

Kyle Gallatin

Yep, they have their own dependencies, of course.

Robby Peralta

And you usually run this stuff on like your own on the same PC. You don't have like uh it's it's not like you have like opsec for you download things here, you scan them, and then you pull them over. This is all just boom, boom, boom, going quickly,

Empathy in ML and Security Collaboration

Speaker 1

right?

Kyle Gallatin

Uh yeah, it depends. But like a big thing with machine learning is experimentation because you're trying a lot of things out. You're not like necessarily you're not always writing source-controlled code that's gonna like end up in a GitHub repository. You might just be like iterating in like a Python notebook, either locally or attached to some like GPU instance in the cloud. So in those instances, like there's not much, you know, like you don't have like automated checks in GitHub for like security vulnerabilities and all that kind of stuff. You're just like installing stuff on some instance, whether it's local or in the cloud, and then experimenting as quickly as possible to try and like get the best model. So there might not be as many like you know scans and things like that as there are on like a nice source controlled repository with like a nice dependency updating and release process on that kind of stuff. Yeah.

Robby Peralta

So um, you know, to sort of like combat those things, I guess it just uh would you agree that it's kind of on sort of your shoulders, like uh as in your teams, to understand the like the threat surface and sort of apply that into your day. Is that is that the best way of going about it? Kind of like software security?

Kyle Gallatin

I think so. I think there's gonna be education that just needs to happen on on both sides. Like what you know, like security folks are gonna have to become more well-versed in the you know, the new attack surfaces and new processes for like ML workflows that introduce risk. And then conversely, of course, they're gonna have to educate the ML practitioners. Like, here's within the confines of what you're, you know, you have to do this, like this is work, yes, where you're gonna be doing. Here's the way that you can do it, you know, safer and better. And here's some tooling that we can provide you to do it safer and better. So I think, you know, the more education there is on both sides, the more we can kind of like work together to achieve achieve the goals. Since we're both here to do the same thing, usually, you know, we're working for a company, trying to make the company succeed.

Robby Peralta

Um, and so how is that education actually happening in reality?

Kyle Gallatin

I feel like the the intersection between like like ML-driven applications and security is is there. Um and it's it might be more about like kind of overcoming that last hurdle of like ML folks being educated and security being educated again rather than like its own field of like talk and study. Because at the end of the end of the day, like ML folks are deploying applications. Like, yes, there's ML, yes, there's like some new attack vectors, but like it's you're trying to protect software and data, which security folks have been doing forever and companies like trying to do forever. Um, and so as long as like security understands like the workflows that ML practitioners need to need to have, for instance, like their data access pattern is going to be different because they need to like access, you know, maybe production data for like experimentation in a way that like typically other workflows might not. And ML practitioners understand that like there are regulations around this data and regulations around the software we deploy, and that they have enough software engineering knowledge not to deploy some kind of general compute engine to the internet that someone can mind Bitcoin on. As long as like each side understands those things, then I think we're I think we're good.

Robby Peralta

Um it kind of sounds like you're saying that like it's kind of like security champions, but just ML champions, or just like a new uh we need a new word for for that, right?

Kyle Gallatin

Yeah, maybe, maybe we do. Is that is that the term security champions?

Robby Peralta

Security champion just means you have like a developer that knows what they're talking about when it comes to security, and they kind of like are helping people around them like get on the bandwagon of like, hey, security wants us to do this, let's do it like this so we satisfy their needs, but also we don't have to stress about changing ourselves too much, right?

Kyle Gallatin

Yeah, I definitely think so. Um, and I would say uh a theme within the ML world is that like I'm a very applied ML person, like very software focused. But there are many ML folks who are, you know, coming from PhD backgrounds who are more theoretical and have less less experience building software applications at scale. So they're used to like working in more sandboxy type environments with kind of like toy data or you know, maybe data for like their postdoc, whatever, um, but not working in like a very like secure setting in the cloud and don't often have as much software engineering experience. So part of my role in the past has been like helping to like educate folks on like, all right, like here are good practices for like developing software with ML. I know you know that. The ML and you know the math, all that kind of stuff. Here's how we build software with it in like an enterprise setting.

Robby Peralta

Yeah. You're like a real, like I'm I'm here for business. I'm trying to make something happen, right? Yeah, exactly. Yeah. Yeah. And uh just while you're there, what are like the nuances there between so like somebody playing with the data and actually making like in a live environment? I guess security is one of them because you actually have people trying to get into your stuff and do nefarious things with it. Is there anything uh else that's notable to mention?

Kyle Gallatin

I would say it just goes back to that like experimentation-focused workflow again, you know, like it's it's interesting, but it's what needs to happen is that ML practitioners, you know, they need to like rapidly try a bunch of things in in a sandbox environment. So they need a Python notebook with access to potentially production data, and they need to install a bunch of stuff in there, and they need to try and run a bunch of stuff on it to see how an ML model might perform. So I think that like that new workflow um is not hasn't like historically been typical, right? There's been nice access control patterns that didn't have to deal with those kind of like these ad hoc experimentation things being run by individuals on certain data sets. And I think that's kind of that's kind of new for some places.

Robby Peralta

I've been asking like my my clients like, do you have somebody that works with AI, ML, like that sort of development team? Do you have somebody? Yeah, do you know them? No. And I was like, okay, why not? Why don't you know them? He's like, yeah, well, they have we haven't had like the reason to, but it sounds like if if if a security team wants to start that conversation, understanding workflow is probably a really good place to start because you just shut up and see how they're doing things today and highlight uh yeah, what what would you what sort of advice would you give to a security team that wants to sneak their way into your world?

Kyle Gallatin

I would yeah, I would say that like definitely training an ML model. Like try doing like a you know, if you there's code labs or quick starts or something for training a model internally at your company, like try those. See what the process says, like um empathy is like the most important thing to have, you know, and like trying to collaborate cross-functionally in any environment. And so understanding the other person's workflow and the things that they'll have to do to achieve their goal is gonna make that conversation a lot smoother when it comes to like, oh, I know you have to do this, so let me provide you with a secure way of doing that kind of thing, as opposed to um, you know, in the past, the reason that they're like I I assume the friction you're referring to with security is that like every now and then security has to come in and say, No, you cannot do this. Um, and that you know, can push timelines, block projects, all that kind of stuff. Um, so if there's an understanding and empathy on on both sides, and folks kind of collaborate from the start and from the outset, like here's what both people understand what the other team is trying to do. Um, I think it's gonna be a more you know beneficial collaboration between those two groups.

Robby Peralta

If only the world had more empathy.

Kyle Gallatin

It's about empathy at the end of the day.

Robby Peralta

It is really uh well, Kyle. Uh I've asked everything that I was wondering about. I've really enjoyed it. Do you have any any last words? Anything that you're looking forward to in the near future?

Kyle Gallatin

Uh I'm looking forward to collaborating more with you know security folks. I've I have a huge respect for and and empathy for security. I also just I've always loved cybersecurity and thought it was really fun. Um, and so I'll just let everyone, every security folk out there know that I'm I'm just I'm trying to be that champion.

Robby Peralta

Um, Kyle, thank you so much. Uh they're not like fixing your house outside, are you? What's what's all the jackhammering going on?

Kyle Gallatin

Dude, they're like just destroying the sidewalk. They're like, and like every corner of of the street is just being like completely destroyed, right? I really hope it's filtered out. Um but uh it's pretty loud. If I was not already awake for this, I would have been woken up by it.

Robby Peralta

Yeah. Yeah, right. Somebody uh somebody with their machine learning model will fix it. Uh somebody from Adobe hopefully will fix it.

Kyle Gallatin

Yeah, I know. I really hope so.

Robby Peralta

I will I will listen to this over again in our second version will be much better than the first because I actually understand a little bit of the life you have.

Kyle Gallatin

So perfectly happy to do it again with another time if we need to.

Robby Peralta

Thank you so much, sir, and uh enjoy your weekend when that time comes. I'm gonna enjoy this Hawaii party.

Kyle Gallatin

Yeah, dude, have fun. Sounds great. Thank you. All right, man.

Robby Peralta

Take care.

Kyle Gallatin

Talk to you later.

Robby Peralta

Well, that's all for today, folks. Thank you for tuning in to the mnemonic security podcast. If you have any concepts or ideas that you'd like us to discuss on future episodes, please feel free to hit me up on LinkedIn or to send us a mail to podcast @ mnemonic.no. Thank you for listening, and we'll see you next time.