mnemonic security podcast

Prompt Engineering

mnemonic

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 26:26

In this episode, Robby welcomes Dan Cleary, Co-founder and CEO of PromptHub, an organisation focused on enhancing LLM-based applications. He’s also one of the organisers behind the Prompt Engineering Conference, the world’s first event exclusively dedicated to prompt engineering, taking place in London on October 16th.

They discuss prompt engineering and management, what to expect from the upcoming conference, and what this emerging field means for enterprise AI and security.

Cleary also shares why centralised prompt management matters, how AI’s role in enterprise adoption is evolving, and what security professionals should explore to stay ahead.

Send us Fan Mail

Speaker

From our headquarters in Oslo, Norway, and on behalf of our host, Robby Peralta, welcome to the mnemonic security podcast.

Robby Peralta

Linguists estimate that there are about 7,000 languages spoken in the world today. Yet large language models don't really understand a single one. So to use them effectively, we've needed to learn a new dialect. Not quite natural language, but not quite code. Prompt engineering. And with this new form of software development, complete with the flourishing open community with countless shared techniques, libraries, and repositories available for everyone, it's destined to be important for security. Which is why I couldn't miss attending this upcoming prompt engineering conference and chatting with one of the organizers to find out what this new discipline really means for our security team. Dan Cleary, welcome to the podcast.

Dan Cleary

Robby, what's up, man? It's good to be chatting with you today.

Robby Peralta

First things first, how was it?

Dan Cleary

Well, if you're referring to the honeymoon and wedding, it all went super well. And yeah, it was beautiful. We're based here in in Manhattan and we went to upstate New York. It was really beautiful weather and everything, no rain, which is basically all you can ask for. I'm definitely still getting used to the whole uh like introducing her as my wife and stuff. But yeah, everything was uh was a ton of fun. Honeymoon was great, but uh happy to be happy to be back.

Robby Peralta

I came across your name through something called the Prompt Engineering Conference, which I'm really looking forward to. And I'm mad at myself for not diving into the topic of prompt engineering before. So what is your background with prompt engineering?

Dan Cleary

Yeah, so I'm one of the co-organizers of the prompt engineering conference in the first time in person this year in London, the previous two years, it was um virtually dude with my fellow organizers, Maxim and Goda. And it's something we kind of just do like on the side because we're generally interested in it. And I run a company that's like in the space called PromptTub. And so we started Prompttub a little shortly after Chat GPT launched. We were working on a different startup at the time. We started to rush to build a bunch of LM-based features into our product and ran into a bunch of issues around testing prompts, versioning them, managing them, um, and you know, essentially built like an internal version of what would become Prompt Hub and thought it was gonna be a much bigger opportunity than what we were working on at the time. And so, yeah, I mean, we've been working on it like for a while now. Um, I've been writing about prompt engineering, you know, basically since since then for over two plus years now. We have one of the most popular blogs on prompt engineering and Substacks. And yeah, it's like a really like interesting topic, and I'm excited to go go deeper on it. Um, and excited to kind of meet and hang out in London in less than two weeks now, which is crazy.

Robby Peralta

Cool. But before Prompt Hub, you were you were in like software development sort of startup vibes.

Dan Cleary

Yeah, I like to say I've never had a real job because I've been running a company my my my whole professional career. Um went to NYU, started like a services company there where we would build apps and websites for people and stuff. Cause that sounded like more fun than trying to get an internship. Um and then yeah, moved to doing like SaaS software stuff, which eventually um led to like a couple of different products. But then yeah, once ChatGPT came out and we saw the kind of opportunity of like new tooling was gonna be really needed for this because it was gonna be a really big thing and it's like completely different than like all other types of like software engineering historically. So there's gonna be a whole new like tool set. Um and we just saw an opportunity there. I mean, it's honestly it's it's either easy to pivot when the previous thing wasn't working super well. So like it was like, okay, this is like you know a little bit of a life raft to a degree. Um and then yeah, we've been kind of off to the races ever since.

Robby Peralta

Yeah, cool. So prompt engineering is semi-important for a person to know. I wish I would have looked into it earlier, but it's really important for like an enterprise use case of LLMs and agents.

Dan Cleary

I think it falls into like two different buckets. So there's like just like the everyday use of like using ChatGPT or some kind of interface like that, where it's just like kind of conversational and like you're just doing it in the browser and stuff. And like in that case, like you don't need to be it's still good to like try and be as specific up front as you can and like do these other things that are just like good best practices in general when using any LLM. It's less important though than if you're gonna be having you know this prompt used in your like application or other internal workflows where it's gonna be run like hundreds of thousands or you know millions of times. And that's in that case, it becomes much more important to like really refine it because at that scale, like little things going wrong will have like a material impact either for you know your team internally or for your users. And so those are the kind of two buckets I like to think about it in. Um, if you're just using it casually, there are some like things that will help you get better outputs and like it will save you time. And I feel like you kind of just pick those up as you just use the tools more. Um, but then yeah, it's much more important when you're actually like shipping LM features, be it internally or externally.

Robby Peralta

So we just had an episode on the topic of like agentic AI, right? And I know you know what that is. So where does prompt engineering fit in agentic AI world?

Dan Cleary

I would say in general, prompts are getting longer, meaning that we're giving the LLMs more things to do and to not do, to specify, to not specify, what tools to use, what tools not to use, the descriptions of the tools, like prompts are generally getting really long. And you could see this through like the companies that have open sourced their like system prompts. And you know, we have a bunch hosted in prompt that you can go check out. But they're they're long. There's sometimes they're like 10,000 tokens and things along those lines. And these are for like agentic applications. And that's just because when you are having these LLMs doing these long running tasks, there's more chance they will run into certain edge cases, and every edge case kind of not every edge case, but the groupings of edge cases need to be accounted for. And a good way to do that is through the through the prompts of the instructions so that you know if the model sees X, you instruct it to do Y. And so the surface area of things that could the model can do is just much larger in like an agentic world, which actually increases the like importance of having really good instructions. And I think, you know, like really descriptive tools is something that we always kind of talk about. You'll see someone put like a lot of work into the system instructions for like a, you know, like a refund agent. And then the like the tool to actually do the refund is just like, you know, the description for that is like really small. Um so that's like a really small like example of how like deep these can can really go and the importance of them um like in this kind of new agentic world.

Robby Peralta

And this agentic world, your platform can you just tell us what it does, uh like the main features and why they're there?

Dan Cleary

Yeah, for sure. So prompt up, the 50 characters or less version is the the GitHub for prompts. And so it's a place where teams can manage, store, version, deploy their prompts. And so it's like the home base for all of like a company's prompts. And there's a whole kind of open source community in the same way that you can open source code on on GitHub and discover it. There's a whole kind of community side to the to the platform as well, where you can see what other prompts people have made public. And we've been lucky to work with with partners like GitHub actually has published a bunch of prompts on our platform. So it's cool to kind of see what uh what other people are doing. But a big benefit that like the companies, enterprises that we work with get is that they're able to store all their prompts in one place. They can make sure that they like work well, they can try different models, they can run like evals and metrics to make sure that things aren't regressing as new models come out. Um it's so it really centralizes everything in one place. And a big thing for us is the focus on uh like a non-technical user. Um, when we first started the company, we saw a lot of tooling out there was like really focused on like a software engineering type of user. And we just knew a lot of people who are gonna be working on prompts and like generally using LLMs or knowing what a good output looks like were gonna be non-technical. And so, yeah, it's basically a prompt library with a bunch of testing features on top. Who uses like these prompts? Is it data engineers or so? A lot of the times it's like, you know, let's take example of like a medical scribe company. So they're building an application that allows doctors to not have to take notes, like the AI will take the notes for them while they're like seeing patients and things along those lines. And so in that case, the users of a tool like ours or like just generally who's working on these LM applications is actually like the medical doctors themselves or medical residents or whatever the different titles may be, but they are the domain experts who know like the medical knowledge. They know from a transcript what a good note would look like from that, and like follow all the right structures and all the right procedures and things along those lines. And then so they're really the ones who are gonna be testing and tweaking and validating and like looking at these things. And then when it comes to like implementation, then you have an engineer or something who will maybe like sync up to our API or like a similar tools API and pull those into production. Um, but the use case is really like those domain experts because they're the ones who know what an output generally looks like. Um, you see a lot of product people just because they're usually pretty close to understanding like the real problems of the uh of the user.

Robby Peralta

So whoever cares about the output is the one that's gonna be tweaking that prompt to make what they want.

Dan Cleary

Yeah, exactly. Because the output is eventually like what is your product that you're selling to your users. Yeah.

Robby Peralta

Yeah. What's that ecosystem look like? Where like when you push out a prompt, where are you pushing it to?

Dan Cleary

Yes. I mean it's pretty simple in terms of like you're refine a prompt in something like our tool, and then you'll usually just like pull that into your product or like into maybe like a workflow builder like Zapier or NADN or some other custom like internal software that you use internally. But you basically just have a API request us to like get the prompt, and then you serve that to your users directly through your code, basically.

Robby Peralta

And how often how do you like measure the efficacy of these these prompts? You just look at it and say, I don't like this, so we change it, or how does that go?

Dan Cleary

Yeah, I mean that you know, the vibe like the vibe check is like I think something that some people think is like not a good thing, but I think it's a very good place to start of just like, okay, yeah, let's just let's go from a starting point. I'm gonna write a first version and I'm just gonna like run it just even a few times and like tweak and run and tweak and run and just use my eyes to see like if we're getting generally to like the right place. Is it getting better, like measurably getting better, just like from my like professional judgment? Um, and so we call that yeah, like the eye test, which is like a very good place to start. Um, and then once you are feeling pretty good about that, you do want to have like some quantitative method, um, which is where something called like evaluations comes into play, which are basically just ways to like systematically and continuously measure your prompts and their outputs. Um, that's consistent. So for example, for that medical scribe example we were talking about before, like maybe the the notes need to have like three headers. One of them needs to be like a conclusion header or something, or like it needs to include the patient details, and the patient details need to have a name, a location, and like the symptoms or something like that. And so that's something you'd use evals for because you want to test it like at scale, because you'll be running them many, many times. So basically just like sets of rules that you can then continuously apply and track not only when you're testing your prompts, but also when your users are actually using them. Um, so you could see if things are you know working in production essentially.

Robby Peralta

So and your platform does this, so you have like some sort of agentic workflows in your product just to help with that, I guess.

Dan Cleary

Yeah, yeah. So we have a bunch of like AI kind of nested into the product in in some ways. We have AI to help you write better prompts. We have AI to like make suggestions based on how these evals are like working and running. Um, and so there's lots of different ways that we also like it's good for us to try and like kind of dog food our product too. But yeah, this you want to have this kind of circular loop of writing a prompt, testing it, running evals, learning from that, making the prompt better, and kind of this continually like improving it. Um, so then your users get a better experience.

Robby Peralta

So how many questions do you get about security in this?

Dan Cleary

Larger companies are very aware and interested about what data they are sending to the model providers because people use this for like a wide range of use cases. People even just use this to like store prompts that their team can then use in like ChatGPT and like just come in and copy it and like be on their way. Um, and so from a security perspective, we are built on like very similar versioning to uh Git and like what GitHub has done for for software. And so the security people actually really like what we're we're doing because it gives like a very clear version history of here's what changed within the prompt, when, who did it. Like there's this full kind of audit trail that's like inherently baked into the versioning of the prompts in the system that we've built. And so you can add rules of like don't accept this prompt if it contains like this language or this or that, um, like these rules, you know, like leaking secrets, things along those lines. And so we have been hearing it more and more recently as like larger companies, I think, are starting to adopt AI like across the board, and they're much more, you know, of course, more security like aware. Um, but I mean it's a huge topic. I'm sure you got you talk about it a lot and hear a lot about just like generally how like enterprises are thinking about like AI and security.

Robby Peralta

We think prompt injection. Al, how many have you gotten to questions just around that? Because I I guess your product could block that per per se. Yeah.

Dan Cleary

So yeah, we definitely help like with that to a degree. And it's actually kind of what first got me into this world because I was it was like I read an interesting article about how you could like embed a prompt into like the source code of like your website, and then you could send the you know the LM to like read that, and then you could like essentially yeah, in prompt injected to like actually do whatever the prompt was on like the underlying website. I was like, Oh, this is actually really cool. That was like two years ago, but yeah, we can help protect against that. And there's like protections you can do like on the there's like multiple layers where you can protect against the prompt injection. One of them is like actually in like the system prompt that you provide. It's not a super robust way to do it, but it will cover like you know, 50% or like some like big, pretty big portion of it.

Robby Peralta

Um tell me how, real quick, that 50%, how that how it actually is.

Dan Cleary

It's even simple stuff of like just what instructions would you want to provide to the model to make sure like it stays on the task, it doesn't deviate from it. It knows that like you know that people might try and persuade it to do otherwise, just like writing instructions that are like specific to your use case of how you think prompt injections might come. And so it's even like really simple language like that can be can be really helpful. And giving the model like an out if it doesn't know what to do is also, I think, generally a good practice, but also helps protect against prompt injections. Um yeah, there's like other steps along the way where you can process the user's input before giving it to the model. Like you can have another LLM check the input real quick before giving it to the model that's actually gonna do something. And so there's other steps along the way where you can try and protect against it as well.

Robby Peralta

So this is like an architecture discussion on how to make sure that these things don't happen. You just put an LLM after an LLM and another LLM and then a prompt at the end of all that.

Dan Cleary

Yeah, you can, you know, and when in doubt, you can throw more intelligence at the problem, and it usually helps. It won't get you all the way. You could also do some like, of course, like some code type checks where you like look through the content for specific things using just like you know general code to do that. But yeah, stacking stacking LMs, small ones, make keep it fast, keep latency down, um, is definitely like a strategy.

Robby Peralta

Where do you think security fits in in this like ecosystem of yours that uh that is that is moving in one direction and it's not gonna go back?

Dan Cleary

Yeah, it's a good question. Um there's a lot of different types of attacks to worry about. MCP kind of opens up this like whole world uh of like different types of attacks that are kind of interesting. Um so I think you want to move quickly. You want to try and find whatever the like 80-20 kind of thing is there where you can implement something small that covers like 80% of the boundaries, and you really want to like just enable your team to do more. I think like there's a lot of things we don't know on the security side, and we're gonna have to kind of figure it out the the hard way and just kind of continue to like empower people, give the tools necessary to do that because it's just like such inevolving space, there's gonna be a lot of like new learnings.

Robby Peralta

Yeah, right. It's going extremely fast, but I really haven't heard of that many like big hacks coming from this side of the fence yet. Have you?

Dan Cleary

Yeah, no, not a lot of like big hacks. There's been like some prop prompt injection things of like I feel like the popular ones are like chatbot. There was one with like an airline, and there was one with like a like a car company where they were able to get like a free car or something. Um I don't even know if you call that like security versus like just like poor for poor performance. Um, but no, there hasn't been a lot of like actual hacks yet, um, at least like big ones.

Robby Peralta

The people listening to this, they're working with cybersecurity. What do you what would you recommend to them to actually like read up on to understand this side of your world and be able to work best with you guys moving forward?

Dan Cleary

Yeah, I mean, I would think start, you know, just to consume as much like content around it as you can. I would definitely say like specifically, like the biggest thing that I would worry about would be on like the MCP side of things, just because I think that opens up the most weirdness that can happen because you have the you have less control if you're connecting to someone else's server to do something or other. Um, of course, you can kind of like roll your own and like do all these other things. But I would suggest having good guardrails in place if you're shipping these things to production, I think is like just very important. And I think the more educating you can do internally about like what those guardrails are, why they are what they are, to like help other people in the company get a better understanding of what the posture is towards these things, I think generally would kind of be a tide that would raise all the boats.

Robby Peralta

You think it's a better idea for somebody that's working more as a developer or comes from like data science and it and AI than a security person? Since the security person just doesn't understand all that.

Dan Cleary

Yeah, I mean, I think generally it's like a good mix for someone like myself as like a general engineer in the area. Like I have a good understanding of how these things work, but a security person would have like a better idea of like what like how those map to like security risks to the business. Um so I think a good mix or like you know, a team working in tandem like that would be kind of the the best of both.

Robby Peralta

Cool. So tell us what's going on in the conference. I said you said you had 25 speakers. What are they gonna be touching on?

Dan Cleary

Yeah, and I think we're even up from that. But yeah, the conference will be really fun. As I mentioned, we've done it for the past two years. This will be the third year, first time in person, which we're really excited about. Um, it's at a cool place in London. And we have a bunch of really good speakers from awesome companies like Box and AWS and um a bunch of like early sh earlier stage startups as well, who are like very much soapboots on the ground. Um and we'll be talking about a lot of different things. We'll be talking about MCP, we'll be talking about agents. We'll of course be talking a lot about prompt engineering, we'll talk about like automated prompt engineering and how LLMs can like sometimes perform better than than humans in that category, no surprise. Um, and kind of what the future of that like generally looks like. Everything that kind of goes into getting better outputs from from LLMs and the different ways that we're using them via agents and stuff. So it should be should be a lot of fun. And we're very much looking forward to it.

Robby Peralta

I literally think that in your future conferences, you could have like a little security uh track of that conference. Uh I think so too. We don't know anything about that space, so we have to go there.

Dan Cleary

Yeah, I think maybe next year um for sure. I think it's like a very popular topic, but I think specifically around guardrails. Just a lot of like information we could we could definitely like dive into.

Robby Peralta

Yeah, because I mean AI, the they're always talking about the audibility of AI. You kind of have to have a platform like yours or full control of the prompts to actually be able to do audible AI or Yeah.

Dan Cleary

I mean, you need to know what's going into the model and what's coming out of it, and having a centralized place for that is is really helpful in like really like a requirement, not only from like a security perspective, but of course from like a product perspective to see how the model's handling certain things in the way that it should be or not.

Robby Peralta

And the last thing you want to pick up on is you said that uh the computers are better at prompting than humans, but that kind of makes sense because we're talking to code uh with our human eyes. We don't understand most people don't understand how the the LMs work. And then an LM knows exactly how an LLM works. Is that actually accurate or is there a misunderstanding there?

Dan Cleary

Um I'd say it's partially accurate. LMs can't do it like all the way, and I would still say, like, generally speaking, you need like a human to kind of to do this to a degree. LMs tend to kind of overfit. So if you're trying to have the LLM up like update a prompt to solve a specific thing, like it will do that, but then it will like overfit to that specific thing and it will kind of try to always like act in that manner and might disregard some of the other like important aspects of the the logic. And so with prompts, like the the challenge really comes down to getting the all of the logic of whatever you're trying to do down on to paper in like a clear and concise way. Um and generally like the LLM won't always know about that logic. And if you've already, if you've gone to the point where the LM knows all the logic, then you've like already done the work. Um so it's really hard for it to intuit what's important. We have a lot of like subconscious understanding of things that we don't always put into the text box. Um that's where like a lot of the challenge comes from, and that's where like LLMs aren't very good as at picking. things up because they don't have like the whole context that we we have that we just like kind of inherently know. Um but there's all these cool like automated frameworks that can like take test cases and rerun them and like there's like a definitely like a place for it and we use it in like some ways but none of them have like really taken off too much.

Robby Peralta

That's the whole point of prompt engineering though is to get that human intuition on paper and into code exactly how we want it. And then you've done your job and you can never thought like that. What makes me think about remember the I think it was like Facebook they made a model where it just started speaking computer language to each other and the humans didn't know what they were talking about anymore.

Dan Cleary

Yeah a lot of weird a lot of weird things at the at the corners and at the edges of all this stuff.

Robby Peralta

Last but not least, what if we talk next year? What are you working with right now? What's what's causing you trouble? And uh what's the thing that you don't have an answer to right now that you hope you solve next year that'll make the world a better place.

Dan Cleary

That's a good question. I hope that we're seeing more like enterprise like org wide adoption. Yeah we see a lot of companies who you know are hitting like the drums to like adopt AI, but I think there's like a big distance between like what the CEO and the founding team wants to do versus like what people are actually like doing. And so I hope that chasm like just gets smaller and smaller where it's more like deeply integrated people are getting better at using it. It's like providing more value to them, which like I think will be of course good for us, but generally good for for everyone as well. And so that's what I hope like in a year from now that basically where people are saying they want to adopt AI and where people are adopting I AI is like much closer.

Robby Peralta

Do you see in your from your side of the fence that companies actually have like an AI center of excellence or one one point for AI that actually works with all the business students to make their dreams and hopes come true?

Dan Cleary

We do see these types of committees popping up and it's like their job to like maybe wrangle up like the heads of all the departments, talk to them, see what they need to do to like implement things to like make their life easier. What are the use cases? What do they need? Um but so we are kind of seeing this you know top-down approach. We also see a lot of like trainings for those people or like outside consultants come in to try and help those people to then yeah work with the other departments, see what their needs are and you know it's it's I mostly the IT because it's like new software is needed to like bring in and change these things. And so it like requires them to be heavily involved in the process. But yeah, we've seen that work and I think it's like a generally a good way to go about things. We've seen a lot of bottoms up approaches too of like okay we're gonna try and like educate everyone in the company as much as we can and empower them versus like waiting for this like committee to come in. But that means you need to like cut the red tape and allow people to actually go and like build stuff and change things and not have rules around it. So we've seen like both approaches work. I think the bottoms up approach is like a little bit more like sustainable and just like faster. But it's like a little scarier I think for for larger companies.

Robby Peralta

Yeah. But I guess you you would love if that actually happened because if you had every division of the company actually using or having its own strategy, its own use cases, then it would more need a a product like yours to help every division manage their their thing in a in a safe way.

Dan Cleary

Yeah. I mean I think just empowerment in general um through our tool through another tool like I think is kind of where you want to be you need people to start like just playing with these things and using them and exploring and like that's like the best way to actually start learning what's possible. And then once you have one use case you discover like five more from that and then um having ways for people to just like actually do that. Because they're being told they're supposed to do it. But a lot of times like they don't actually have the tools they need to actually do it.

Robby Peralta

Yeah. They want to start doing it but they have to do it in their in the in their evenings because they haven't gotten like a mandate to use X percentage of their time to actually go and go out of their way and do that because they have to be doing something else right. Yeah but Mr. Cleary uh I'm really looking forward to seeing you in London beer on me beers. Yeah it's gonna be a lot of fun they call them pints over there no they they call them pints yeah I don't really know what that means I just know it's bigger than the beer I'm used to so that's not a bad thing. I like pints.

Dan Cleary

Right, right.

Robby Peralta

So uh cool thank you so much for time Mr. Cleary and I will see you here in a few weeks in London.

Dan Cleary

Sounds good Robby thanks

Robby Peralta

until then ciao well that's all for today folks thank you for tuning in to the mnemonic security podcast if you have any concepts or ideas that you'd like us to discuss on future episodes please feel free to hit me up on LinkedIn or to send us the mail to podcast at mnemonic.no thank you for listening and we'll see you next time