Build or Buy AI Pentesting? What the POC Doesn't Tell You
Watch Synack customer Dow discuss why a successful AI pentesting POC is only the beginning, and what it takes to deliver reliable, validated security outcomes at enterprise scale.
Overview
Watch Synack customer Dow discuss why a successful AI pentesting POC is only the beginning, and what it takes to deliver reliable, validated security outcomes at enterprise scale.
What you'll learn
Why a working POC is not the same as a dependable security control
A proof of concept shows an agent can find something. It does not show the capability will hold up as a control the business depends on, across a changing attack surface.
The hidden costs that can derail internal builds
Token spend, model drift, staffing and compliance gaps. Dan Lacher puts it plainly at 6:36: it is a lot of tokens and a lot more engineering manpower to rebuild it all from scratch.
Why false positives are an architectural problem
Noisy output is a property of how the system decides something is exploitable, not a threshold to be tuned. Mark Kuhr traces it back to the agent harness, not the model.
What human oversight looks like in practice
Humans direct the agents before testing and validate what comes back. Dow does not pass agent output to its support teams until a person has confirmed what is genuinely exploitable in their environment.
The questions to ask before committing engineering resources
What the harness has to do, who maintains it as models change, how it is benchmarked against real enterprise apps rather than labs, and who is staffed to run it long term.
Chapters
Jump straight to the part you need.
Full transcript
Read transcript
Good afternoon, everyone. Thanks for jumping on today's webinar. My name is Tim Norgavitt. I am a senior manager of the solutions architect team here at SYNNAC, and we've got a fun conversation lined up for today. We're we're really gonna dig into this whole buy versus build topic. We're gonna, you know, start with Dan, talk about how Dow ended up choosing SYNNAC for penetration testing, kinda how they went the buy versus build route, and then we're gonna flip it over and kinda talk with Mark about how Synack built its own AI pen testing engine, or Synack Autonomous Red Agent, or SAR as most of us call it.
But before we dive in, quick housekeeping. If you have any questions today, feel free to drop them in the chat, and we'll do our best to circle back as to as many of them as we can, and maybe follow-up afterwards if needed. But before we get anything, let's introduce our speakers today. Dan, you wanna kick things off and give everyone a quick intro? Sure. Dan Locker at Dow, cyber engineering team leader here. I've been here for seventeen years and eight years in the cyber area.
Cool. Thanks, Dan. Mark, over to you. What do you do, and who are you? I'm Mark Core, cofounder and CTO here at Cenac. Run all things product and engineering here. I've been pretty involved in a lot of our AI build out and initiatives with Sarah, as you mentioned. I I do refer to her as Sarah. I give her, I guess, a name there that would be familiar to humans even though we we do see it as an acronym as well. Thanks, Mark. Yeah. Looking forward to digging into that in a little bit.
So, again, my name's Tim. I'll be steering the ship today as we kind of have this conversation, try to get both sides of this buy versus build debate. But Dan, I'd like to start the conversation with you. So for a minute, I'd to go back to two thousand twenty two, that's kind of when we started this partnership with Dow and Synack. Kind of a multi part question to kick things off, know, what was going on several years ago at that time? Kind of what problem were you trying to solve, and maybe what else, you know, what other alternatives did you look at before you landed on
pursuing this partnership with Synak? Yeah. We had had our own internal red team with starts and stops as people had exited the company or transitioned roles, and we're really looking for a value solution to continue to give us compliance and really to start to go after some of our highest value web facing assets. As we started to look through those, we ended up with SYNNAC.
Have had you guys on board since twenty twenty two, like you said. And at that time, we didn't have an internal red team, and so we really had to build an engagement with someone that we could focus on, and we were really looking for the continuous coverage options that Cenac provided us. The the point in time, you know, we'd get the results and then turn a blind eye to it for several months and then come back at it, you know, maybe once a year, and it was just something we really wanted to close the gap
on with the continuous coverage option. That's what we're hearing a lot, kind of the point in time transition to a more continuous ongoing assessment. The point in time, the second you get the report, it's out of date at that point in time. So you mentioned the red team kinda had you had some starts and stops, and there wasn't, like, an official red team when we started. Kind of fast forward today, like, how has that changed and evolved over time, and what does that red team look like today? Yeah. We've we did we have rebuilt our internal team.
It's small at this point, but we are using them for a lot of internally focused engagements and then building out structured engagements to hand to SYNNAC, as well as our internal teams are the only ones that know what assets we need to be going after. So it's really helping drive and direct where we're putting the SYNNAC power to externally, as well as soon to be internally on our network.
That's great to hear. So if I can make a kind of follow-up on that, it sounds like buying and kind of partnering with SynAct didn't necessarily sink the need for internal red team, and actually kind of leveled it off. Would you say that's correct? Yeah. I I don't know who gave me this term, but I really see it as a force multiplier, right? I have a small team, but when I can engage with SYNNAC, I have the the power almost unlimited scale with SYNNAC to, you know, uses that force multiplier for my team to just go across assets at scale and speed and velocity.
Yeah. So kinda digging a little more today and kind of aligning with the conversation we were to talk about, you know, let's look at AI pen tests. I know that caught your your in your team's eye at some point. I I'd like to dig into there a little bit that I I know from talking to some of the people on the account team at Sync here that supports you all, your goal was never to build, like, a full end to end AI pen testing solution. So talk us a little through that thought process. Kinda did you look at other AI pen testing vendors out there? Kind of how did you decide on what you're gonna own versus
outsource, kind of where did you end up with things, and what do you what does that look like today? Yeah. There's there's a lot of there's a lot of them on the market. We've looked at a few, did POCs and proof of values with a few of them, brought in these tools that are really, you know, turn them loose, they go off, they assess, a few of the the better ones are starting to learn to chain results together and then kick out a report at the end.
Again, those tools give us that sort of full force multiplier effect because we have to have our humans directing where we put those tools, and then, you know, digesting the results afterwards. And, like I said, looked at a few of them, didn't go with them, and then now that we've restarted the red team, it we've looked at what it would take to just build that all ourself. It's a lot of tokens and and a lot more engineering
manpower to rebuild that all from scratch versus have an entire engineering arm at a trusted vendor that builds it and we can use the output from there. I think it's a really interesting application of of kind of humans and AI working together on those engagements as well. Finding ways where, you know, you deploy the the AI agents to go, you know, maybe do things that are broader,
maybe not as not as in-depth, and then you use the humans to guide and and do follow on testing with with things and go a little bit deeper. Is that is that kinda how you're seeing, you know, humans and AI working together, Dan? Yeah. It just gives us that ability to again, we have to have the humans directing them to start with. The the agents never sleep, so we can get the scale and the speed at that, but the results coming out of that,
we can't just kick those off to our support teams right away. We gotta put that human back in, look at the results, see what actually is meaningful, and every one of these AI platforms for pen testing don't know the additional defense and depth layers we have of, right, there might be a vulnerability there, but is it really exploitable or do we have other mitigating factors there? And that that human that knows that is very helpful. Yeah. Think that that guidance and that context that the the the
internal team can give the AI agents just super crucial in the way it evaluates potential vulnerabilities and and also the way it hunts. So I think that that context mismatch where you have if you have no AI with AI without context, it's it's gonna get lost in the in the sauce a bit. And so I think it's important that you you have that internal documentation. And and then the same way you would onboard a human pen test team is the same way you wanna onboard an agent pen test team.
Kinda building on that a little bit, know, Dan, kind of what you and your team are building, from what I gather, is that you're building something kinda focused on like the external attack surface, so you have better data to feed into the testing, is that correct? Yeah, we've really taken on the SYNNAC in multiple pieces that is offered there. One, you know, initially, we just put the continuous testing against our highest valued external assets,
ensure we had coverage, ensure we can find and fix at speed and scale, but we have immense breadth and the breadth that's facing externally grows at a rapid pace. So what we've started to do is then take our internal teams and through automation, look at those things that we expose and hand them off to SYNNAC in near real time so that we can turn around and have SYNNAC come back at those
assets with the scale that SYNNAC has for that. Right? You could think of engineers exposing an s three bucket or blob storage, and as new assets are delivered to the Internet, the ability to scrape that out of our cloud and feed it back to SYNNAC at near real time so that we can you can turn around and give us the results back on the asset or or where we've built some automations with our
internal team to do. Yeah. That's pretty cool. That's a really great way to use the Synapse platform and pushing data to it, getting the testing exactly what you need based upon the the attack surface. That's a that's a really intelligent way to apply the tech. You know, Mark, maybe you can add a comment here because I think this speaks to why we built SARA in the first place. It wasn't a matter of trying to build the shiny new object and just have a cool new tool. What Dan just spoke to is a problem we see with a lot of
customers where the attack surface is growing faster than most security teams can keep up with. Yeah. It's a com it's a common problem. I mean, you try to attest this to human pentest teams with all these changes, they're gonna drown in all the changes. They're gonna drown in the number of vaults. And so if you hand those to AI agents and and especially ones that are built to scale out and and into teams, you're gonna end up with, you know, a lot of lot of horsepower you didn't have before to go dig into those issues, figure out which ones are actually exploitable,
suppress the ones that aren't, and then have your your human team, your remediation efforts focused on the actual issues that adversaries could exploit. And this is how we clean that that signal up and get rid of the noise. And it's just really critical that the that we have a platform where that humans and AI agents can collaborate together, sharing information back and forth, sharing context, really move the ball forward where, you know, these agents can be, you know, supporting actors to the human teams and and really bolster
their capacity. Absolutely. Yeah. So one last question for you, Dan, before we kind of dig in with Mark and kind of how we built Sarah. So when you look back at twenty twenty two and maybe even now kind of partner with us with with Sarah, like, what was it that about SynAct that kind of clicked as where someone that, you know, you can move forward with? Like, what mattered most to you about the team about whether I test on our platform? Maybe speak to that for a minute. Yeah.
I mean, it comes back to, and this is what brought us to you guys in the first place, well established, vetted researchers was the key piece there. As we were going to give you access to our most valuable assets, we we needed the security research to be very well vetted. A company that is backed and reliable and then as we moved into Sarah,
it's the human in the loop aspect and I'm sure Mark is gonna touch on this, that even after the AI goes out and looks at it, finds it, tries to exploit it, before I see it then as a customer and I get the reports back, I know that a human is vetted exactly what the AI has found and it's at that same level that I get from human validated testing as well. Yeah. Yeah. You don't wanna waste your time chasing, you know,
AI generated swap and false positives. I think that's that's where people are getting a bit lost now is putting agents on things without any human in the loop can can be quite noisy. So it's it's good to see it all come together in into one flow where you've got this quality verification and, frankly, a learning loop where we're getting smarter and smarter on on making the AI agents better at discerning exploitable wounds from ones that aren't. And then more specifically, you know,
getting to the point where your own acceptance rules, you know, for your business, you may find certain things in certain areas of the network. It's an acceptable risk, and and that's okay. And so that that kind of context on your business and the way you run it and whatnot needs to come into the agent decision making process. And I think that's a great setup to kind of transition and dig a little more into Sarah Mark and take us through the build process. Know, When we talk to people about AgenTic AI pen tests and their
mind jumps straight to the models. So Yeah. May maybe give us, you know, your perspective on what when you looked at the models for Sarah, kind of what your thought process was and maybe what some of the trade offs were and explain some of that. Yeah. I mean, the models the models do a lot of the thinking and the reasoning in these systems, but they're not they're not the whole picture. You know, a lot of this, you know, delivery excellence comes from the harness that you put around the models. And that's that's really, really key. Like and we've run experiments where we've taken the base
models from the large frontier labs and also the open models that are that are floating around and and just said, hey. Here's a prompt. Run a pen test. Do this thing against this lab. And then compare that to, you know, the the harness wrapped model, and the performance isn't even close. It you know, the harness harness really makes a big difference because it guides the model's decision making criteria. Whereas the models are getting smarter and smarter, and I think we we will see that over time, the harness is still very critical.
And and you can see this in other areas too. If you look at software engineering, for example, and follow developers who are using this agentic coding for for their daily tasks, just not not having the right context for the problem creates really sloppy code. Not having the right, you know, rules for the agent on how it writes code also creates some slop. So it's really important in in any of these domains where we're applying these NGENTIX systems and these models to get the harness right for your particular use case.
And SYNNAC has done that with Sarah by, you know, really learning on on all the pen tests that we do, seeing how it performs, seeing would it make the same decisions that that we would make about a pen test, about a particular vault, about a particular attack surface, about a particular, you know you know, type of fingerprint that we're seeing in the scan. You know? So we're we're really digging in on the particulars of how a human pen tester would go after a target and bringing that into the harness.
And so it gets better and better over time, and that that isn't requiring anything at the model layer. You can extend the models, and you can make the case that, yes, the future will be, you know, custom models, and we're gonna build our own models. And we have experiments in that in that vein as well. But that that's a very large engineering effort in itself to to code your own model up and and and get that prepared. And so I think this build versus buy discussion, people quickly get to a level of technical complexity where
they're just not comfortable. And it's like, wow, I I really bit off more than I can chew with this. I'd rather just outsource the building of this to somebody else and let me benefit from from the things that that my business needs, which is a more scalable solution for pen testing or a more scalable way to go through all the vulnerability findings from the scanners. Whatever the the problem is, you know, let your business focus on that problem and and procure a solution to to fit that. Of course, I'm biased in that because we make one. But the point is,
it's more difficult than people realize to build one of these agentic pen test capabilities. And and you can try it. Go go into Clogcode and say, build me a pen testing agent and framework, and and it will build you something. It will not work the same way a human pen testing works. Promise you that. So you referenced experiments in there a couple times and kind of monitoring some of the early models. Like, everyone's favorite topic when you talk about this is benchmarking. So maybe take us back to how you first looked at
benchmarking with Sarah and some of the early labs and why maybe that approach ended up being a trap. Yeah. Yeah. The labs the labs are interesting, and and you can get great amazing results in labs. And a lot of the labs are designed to be simple capture the flag exercises of, you know, I'm a build a small app. It's gonna possess this particular cross site scripting vulnerability, and it's and and the the model must, you know, navigate through the pages and send a particular payload to exploit it.
And then you capture the flag. There's lots of those floating around. Unfortunately, those types of labs, they don't they don't represent the complexity of the real world. And so when you when you compare how your agent performs the labs versus real world against enterprise apps that have real logins, real complicated workflows, a lot of graphics, you know, different things like that, it's it's it's not gonna be the same. And so it really breaks down. If you built based on lab environment,
when you get to the real world, it's not gonna perform the same way. And so that's a that's a trap a lot of people fall into. And so, you know, this is an area where I think we have a we have an advantage as a as a builder, because we see lots of different types of apps, and we see lots of different complexities out there in the enterprise space, whether it be network appliances and devices that behave differently based upon their fingerprint, based upon their OS version, or we see apps that, you know, are frankly just just all sorts of varieties of apps,
all sorts of different back end web technologies, versions of JavaScript, you name it. The complexity and the randomness is quite huge in the app space. That's that's where you really gotta cut your teeth when you're building the harness is get that exposure and bring those learnings back into the harness. Yeah. You you mentioned kind of the randomness there, and kinda, Dan, back to you. You kind of talked about that freshness gap earlier, and how you're working on making sure the platform near real time knows what assets show up.
You know, does that sound kinda like the the real world messiness that Mark's talking about is what's showing up in the labs and not mapped from production? Yeah, absolutely. I mean, we have lots of smart engineers that are doing amazing things here, and you know, the speed at which they want to move and again, even internally with AI, the speed at which our engineers are moving, it's just that ability to cut down on and catching it earlier before we
get it to production is, you know, where we're using it. And I don't want to I don't wanna turn my a or red team into AI masters, right? I want we don't need to go build our own models and to Mark's point, we don't need to every two, three months figure out what new frontier model is going to accelerate us faster. You know, it's fun to go do that in my home lab, but I don't need to be doing that at production and enterprise.
And it's kind of back to what Mark said, know, do what you're good at. You know, Dow, you are a chemical company. Focus on that and securing that and kinda partner with someone like Synak that we have our expertise as well. Yeah. So, Mark, you kinda talked about the difference between Sarah in the labs and then actually once we got out in the real world in production, maybe talk about what it actually took kinda making that transition and making Sarah ready for, like, actual customer environment service production. Like, what type of engineering work?
Like, what did you have to do? Maybe lessons learned from that. Yeah. I think one of the things that that we learned quickly is that the guardrails process needs to be fairly rigorous. You know, there's lots of stories on the Internet now of agents making strange decisions to, you know, delete random directories and delete files in production and and, you know, the latest one with OpenAI and Hugging Face and the model deciding it should it should break out of its sandbox to get the key to the lab and lab tests.
You know? And one of the things that we ended up doing with guardrails is is adding a a classifier at the time of execution. And this is not something we we needed really in the lab environment, but when we got to production, there were so many different avenues to go down and so many different attack paths to try. The creativity of the models, while while not as creative as a human, they're still pretty creative on how they wanna circumvent and achieve the goals that you give them. And so we ended up adding a pretty rigorous guardrails
stack where everything has to run through a a specialized risk classifier that makes sure that we're adhering to the rules of engagement set forth in in the pen test. And in a in a pen test platform like ours, every customer can kinda customize those things. You've got you've got scope rules. You've got, you know, things, maybe areas of your app you don't want to focus on or you don't wanna follow certain flows and you're like, hey, ignore this area in production or or whatnot. All those rules have to be enforced at runtime across a
swarm of agents. And so our our guardrail system handles that by classifying the commands in real time really quickly and and allowing the agents to to continue mission while also, you know, making sure it's it's production safe. So that was a that was a really fun engineering project. You you mentioned kind of the swarm of agents. I've read this multiple times that this is kind of the direction that the industry is moving. Like, I read an anthropic report where on their own end,
they set up compared one kind of lead, fully autonomous agent against a bunch of smaller sub agents led by one orchestrating agent, and the Swarm agents beat them ninety percent of the time on the internal benchmarks. And Microsoft, same thing. They built a system with over one hundred specialized agents and found that it beat both Anthropic and Open Eyes frontier models on all their cybersecurity benchmarks with the swarm approach versus a single. Yeah. And and I think that's a good approach, you know, to take.
You know? I think we're we're moving to a place where where you're gonna end up on the continuous tests, you know, like like Dan was talking about before. You know, the continuous is is the best product, and and Cenac has offered that, you know, for for thirteen years that we've been in business. And and it is continuous humans for a year on your apps or your infrastructure. And then we rotate those groups of people about every six to eight weeks. The system will rotate, and you'll get new testers on those. And I think that's that's the best way to do security.
It has been for the last decade. I think it's pretty rigorous. It's probably the most most comprehensive continuous test you can buy on the market. And the way that evolves with agents is is really it's gonna be about the same. You're gonna end up with with swarms of agents and swarms of humans. And the swarms of agents are gonna be specialized for their attack surface specialties, and they may use different models. So, you know, what we're doing is looking at benchmarks where, know, we may we may do reconnaissance across the target with the with
with with one agent and then start to break out specialized attack agents to go after particular things of interest. And those particular attack agents will have skills associated with them of, like, some of the best hackers in the world. And so they're they're gonna be really knowledgeable about that particular attack vector. And then they're gonna be able to go deep with fresh context. And this is where it's starting to touch, you know, context management and how you engineer these things because as you as you run tests, you're looking at lots of data,
you're filling up those context windows. And even the the top tier models only have about a million context window size. And we'll get to two, we'll get to three, but it's it's you get drastically different performance with fewer fewer things in the context window. So I think it's really important to manage that along with the sub agents really well and fan out and and slim down where you need to. But that's really the the magic of the harness and and how it controls for that so you get the best performance with the best model, the best harness,
the best specialized agent on a particular piece of the attack surface at any given time. And, course, it helps with speed. If you need to execute across a large attack surface, having lots of agents prosecuting the target at the same time is is how that's done. And this is frankly how adversaries will evolve. And I think the pivot conversation for this is that the defender has to now deal with, you know, these relatively smart specialized attack agents coming all at the same time. And I think that changes the defender game quite a bit.
And I wanted to get a little bit of of insight from Dan as well on how he's thinking about that for his team. Yeah. That's the interesting side is I also have most of our engineering on the defense piece, and it's how can we use AI to defend against AI? Yeah. You know, so we're we have pieces doing that for our own internal use of AI and making sure we
have security layers and boundaries on, you know, all of our employees and our engineers using AI at scale. But then how do we get that for the to respond at speed? So we are the SOC portion of us of Dow that does response and our detection engineering is really looking at building a lot of autonomous features there to respond at scale and at
speed because it takes a while to page somebody out and get them up and get them back on a console where we need to be able to respond to those adversarial threats at machine speed and then across the scale of our global enterprise. Yeah. It really changes the game a lot. Right? And it you know, this this speed of execution on the on the offense is gonna lead to speed of execution on the defense. And I think you quickly get to a place where those initial
triage decisions and maybe initial defensive actions are fully automated because there'll be no other way to move as fast as the adversary if we don't. Yeah. The other part I think there that we're looking at is as mythos, you know, came out and as we're looking at these frontier models, and they're finding new voles that have lived there for, you know, multiple decades, and everyone is gonna have access to these
soon, if they don't already. Sure. We we can't patch our way past the speed of AI, so we have to come up with the new ways of building the defense at scale and speed because, you know, I can't, as new vulnerabilities come out on a a regular basis, we already struggle sometimes to patch on a monthly cadence. There's no way we can start to patch on daily basis. Yeah.
And at the same time, it it gives you some opportunity, though, to to take a fresh look at tech debt that maybe didn't have capacity to take on in the past with with the genetic engineering being a way to to accelerate our our ability to make changes. You know, as we talk about the the speed of the adversary and how they're they're powered by AI, I think it even digs more into the argument that point in time testing is no longer adequate. Like, we we've got to stay we've got to be as fast or faster than the adversary.
So, Dan, I just wanna kinda go back to you again. You know, you talked about how Dow wanted to transition from this point in time to continuous tests, and this was starting before AI was even kind of a a factor in the equation. So what was the decision back then to start looking at continuous testing and why did Dow want to move in that direction? The continuous testing was just the new way that we wanted to start to look at it. We knew that the point in time gave us just that.
That the moment we ended the test and started to generate a report, the data was already stale. The environment continued to change, Our developers, especially for our our main ecommerce websites, were continuing to push out new code, and if we only looked at that, you know, level of efficiency once a year, once a quarter, Right? We just left the door open to new vulnerabilities for a long time.
You you're not gonna leave your front door wide open all night to so we just wanted that continuous evolution on on our most valuable assets. That makes a lot of sense. Thanks for that that follow-up. So, Mark, I I got one more question for you before we start to wrap things up. Kinda Dan touched on this earlier with the token cost and the engineering cost that they wanted to avoid by building their own solution. For customers that partner with SYNNAC and kinda utilize us for
our agentic AI solution, you know, what costs are they dodging by not having to build this in house? Well, we've gone to a model that's fixed price. So, you know, you've got one price to do a test regardless of how many tokens are used on the back end, and it's it's your your guess is as good as mine as on how many tokens actually are gonna get consumed in these tests. I mean, we can estimate, but each run is a little bit different, and each attack surface is a bit different. So what we've done is we've we've simplified the business
model so you don't have to think about, well, is this is this a small amount or a large amount of tokens? Like, that's not a decision that that most people wanna be have to make when they wanna do a test. I just wanna say, hey. I've I've got this app. I wanna test it, and let's let's get it done. And and we give you a simple two tier approach, Serapentest and Serapentest Plus, where there's different scope definitions, but it's a simple fixed price to do the test, and we just simplify that whole decision making process.
And beyond the token cost, I mean, is there an engineering cost customers need to think about when you think about model deprecation and having to retune as models transition? I mean, yeah, if you're gonna build one of these yourself, I mean, there's a lot that goes into making sure it doesn't go out of date quickly. With each model change, you've gotta evaluate your prompts. With each tool update that's in your toolkit, you've gotta make sure it doesn't break anything. And there's a ton of engineering that goes into the swarms and and how the agents interact at the at the time of execution as well. How do you orchestrate that?
How do you make sure it doesn't get out of control? How do you make sure it shuts down properly? There's all these these problems that come up as you're as you're thinking through how to actually execute that kind of pentest that scale when you start scaling out across lots of infrastructure to do it. And then there's also the the thing you gotta worry about, which is, you know, are you getting the best results from any particular model and how do you know? At least with with SYNACK, we can see performance compared to what we'd see in human
performance and see how close the agents are getting to human performance on a variety of of measures. And we have human review on the on the detections so we can also start to feedback learnings on false positives as well. Yeah. I think that kinda aligns with what Dan was talking about earlier that, you know, obviously, the token cost, that that's the cost everyone jumps to, but then there's the staffing. It's one thing to build it, a proof of concept, but actually make it production ready. It's a lot a lot of the customers I talk to,
they don't think through that that you have to staff and manage it long term. So I think we can start wrapping things up here. Know, I think, Dan, between you talking about your journey with Cenac and kind of what you're working on and Mark talking about what we built here at Cenac, I think we've hit a lot of the key points customers and organizations you need to think about in this buy versus build kind of decision. So as we close out here, I'm gonna ask you each a similar question. Dan, I'll start with you. But if there's another security leader that came to you wrestling with this exact same buy versus build call,
kind of what would you tell them? What advice would you pass along to them? I would tell them to buy. On the things you just hit, the engineering effort to have AI data scientists and the continuous engineering effort of just building the scaffolding and the latest models, let let somebody else that has the engineering capacity do that and turn my red team loose against the things that
are are really valuable to us. We know our infrastructure. We know our assets. Let us direct let our internal team direct, you know, someone like Cenac and Sarah at the assets that most valuable to us. It's I I go back to the what I've said earlier. It's a force multiplier for my team. Right? A small team that just has the entire context behind it to have this continuous rotating set
of highly vetted researchers on target for us. I definitely appreciate that insight, Dan. So Mark, same question to you. If a security leader came to you wanting to build a solution themselves, what would you tell them? Well, I think it's a it's an admirable thing to do, but you gotta make sure you're staffed up to do it, and you're prepared. And some some teams have, you know, substantial engineering resources and and budget for doing that. And I think that's a that's a journey you can choose to go down.
But, you know, like for most things, you know, where it's not core to our business, I think outsourcing is a is a better better approach. And, you know, for for us, especially at at Cynack, I mean, when we we think about building an agent like this, it's really about, you know, trying to take every learning that we can we can glean from the thousands of pentests that we do in a year and bring that into our harness as a way to to make the agents more effective. It's really important that we have, you know,
the ability to spot issues in the runtime at scale. We can fan them fan out agents reliably, and we can deliver production safe agent pentas that that is really effective and complements the human tester executions at the same time. Because I think the future of this is gonna be man and machine working together, agent teams working with human teams to make continuous testing at scale just the normal way of doing business.
No. Absolutely. You know, it reminds me of a recent Forrester report I read where they predicted that companies are gonna hold back. The average company is gonna roll back a quarter of their AI spend in twenty twenty seven once this hype settles and the bills start coming in with the tokens and the maintenance costs and everything. I I think the the market is aligning with everything you both just said, that using AI to augment, not replace, your current program. So, Dan, Mark, yeah, thank you both for your time.
Like, seriously, I think it's pretty rare to get an actual customer kind of live in this decision right now, and a CTO who's walked the build path kind of in the same table kind of having this live conversation. Before we wrap up and kind of everyone hops off, you know, if today's conversation got you thinking about where AI can fit into your overall program, We are running a free Sarah pen test right now. We're gonna throw up a slide on the screen that gives you a URL you can go to and request that. So feel free to grab that and reach out to us, and we can talk about how AI fits into your overall program.
But I appreciate everyone's time today. Thanks for joining our webinar and thanks for hanging out with us.
No lines match that search.
Speakers
Dow
Cybersecurity Engineering Team Leader
CTO and Co-founder of Synack
Senior Manager, North America Solutions Architecture
AI pentesting build versus buy FAQ
What is the build versus buy debate in AI pentesting?
Why did Dow choose to partner with Synack instead of building its own AI pentesting tool?
What role do human researchers play alongside Synack's AI agent, Sara?
Why don't AI pentesting benchmarks and lab tests reflect real-world performance?
What hidden costs come with building an in-house AI pentesting engine?
What is Sara, the Synack Autonomous Red Agent?
Keep exploring
Work through the build or buy decision
Build or Buy AI Pentesting? What the POC Doesn't Tell You
39 min
Build vs. Buy for AI Pentesting: Top 5 Questions to Ask
The written companion the customer page itself links to.
Read the Blog →How Accenture Turned Pentesting Into a Force Multiplier
The same argument from a second enterprise security team.
Read The Story →Sara AI Pentesting in Action
Interactive walkthrough of discovery, validation and reporting.
See Sara AI →Free Sara AI Pentest
Run a real AI-powered pentest on an approved target and see what gets validated.
Apply Now →Watch next
41 min Oct 8, 2026 Measuring Risk and Business Impact: A CISO's Approach to Decreasing MTTR Kris Burkhardt, CISO at Accenture, and Rob Cross at Synack on how enterprise security leaders measure cyber risk, reduce MTTR, and connect security outcomes to business… Kris Burkhardt Accenture 41 min watch
Demo Series 11 min Nov 25, 2025 Synack's Agentic AI Pentest for Speed and Scale This video is a practical walkthrough of how Synack’s Agentic AI delivers fast, scalable, and high-certainty pentesting. Learn how security teams can offload high-volume compliance testing… Mark Kuhr Synack 11 min watch
Unplugged 4 min Oct 3, 2025 Pentesting for Compliance + Risk Reduction Learn how Synack pentesting can bridge the gap between your compliance floor and your risk reduction ceiling. Melissa Wooten Synack 4 min watch Next step
Find out what an AI pentest returns before you build one.
Run a Sara AI pentest against a defined scope and see which findings the Synack Red Team confirms as exploitable. It is a faster answer to the build versus buy question than a proof of concept.


