26 minOct 9, 2026

AI Pentesting in Practice: How Paramount Expands Security Coverage

Paramount's lead information security analyst on moving from periodic testing to continuous coverage, why validation is what keeps more coverage from becoming more noise, and what the Sara harness is doing underneath.

Jessica Boy Lead Information Security Analyst, Paramount
James Duggan Solutions Architect, UK & Ireland

Overview

Paramount’s lead information security analyst on moving from periodic testing to continuous coverage, why validation is what keeps more coverage from becoming more noise, and what the Sara harness is doing underneath.

What you'll learn

Expand coverage

Test more of the attack surface instead of relying on limited, point-in-time pentests.

Focus on what matters

Watch AI pentesting find and validate vulnerabilities, stage by stage, from reconnaissance to a human-validated report.

See Sara in action

Watch AI pentesting find and validate vulnerabilities, stage by stage, from reconnaissance to a human-validated report.

Chapters

Jump straight to the part you need.

Full transcript

Read transcript

Hi, everyone. Welcome to the AI pen testing webinar. If you have any questions during the session, please enter those into the chat. We will be monitoring and answering them throughout the webinar. If we do not get to your question, we will reach out post webinar to answer your question. I'm introducing James Duggan, our SYNNEC representative. Hi, everyone. So I'm James Duggan, as Kim just mentioned. So I'm a solutions architect here at SYNNAC. So I'm based over in the UK, with a background in defense and security,

having delivered various cyber capabilities. And now I help customers improve how they approach their security testing. But today, I'm excited to be here with Jessica Boy, lead information security analyst at Paramount. So, Jessica, would you please introduce yourself and give us a broad overview of your journey in the tech industry and how you came to be at Paramount? Hi, James. Thank you so much for that. My name is Jessica Boy. I am the lead information security analyst at Paramount. I've been in the cybersecurity sector for ten plus years

focusing in application security. My role in Paramount has always been a focus on strategizing application security, make sure we are enforcing, any attack serve as services and making sure there is proper coverage for the company overall. Thanks, Jessica. Really interesting. So just a few questions for you. So with the introduction of AI, how has this changed for you what good security coverage looks like?

AI changed the position of good enough coverage. In the past, security testing was more periodic and focused on known critical assets. Today, applications, APIs, cloud services, and the AI integrations evolve too fast for snapshot based testing. From an AppSec perspective, good coverage now means continuous visibility across much larger attack surface, including shadow APIs, AI enabled features, and rapidly changing environments.

The expectation is no longer did we test it, but how quickly can we identify and validate risk as the environment changes. Completely agree. And knowing that AI powered attackers are finding and exploiting vulnerabilities faster than ever, with the time to exploit window collapsing, how is your infosec strategy changing with that? Attackers are accelerating discovery and exploitation with AI, so defenders have to accelerate validation and response.

For us, that means shifting from reactive security to continuous testing and prior prioritization. We focus heavily on signal quality, exploitability, and reducing the time between exposure and remediation. The strategy is less about adding more alerts, but more about creating faster feedback loops between security, engineering, and operations. Thanks. So just kind of on that, and you mentioned kind of, a bit about kind of faster feedback loops.

Thinking about kind of the more kind of the emergence of kind of the latest AI models, particularly frontier AI models, so the likes of, GPT five point five, MEFOS, which I'm sure all of us have heard about. What's what's your thoughts on kind of using those frontier models either for pentesting or for other security functions? We see frontier models as four smart suppliers, not replacements. They can dramatically improve scale, coverage, and speed in areas like reconnaissance,

pattern analysis, and identify potential attack paths. But in enterprise environments, context matters. Human expertise is still critical for validating business logic issues, understanding risk impact, and separating noise from meaningful findings. The real value comes from combining AI driven efficiency with experienced security practitioners who understand the environment and the business. So, yeah, that's really interesting.

So do you think, from from your view, humans are certainly gonna remain a critical part within within that cycle of testing? It's not gonna be purely AI. So do you how do you see the sort of blend of where sort of humans might sort of remain to be sort of the critical part? Absolutely, James. The strongest model is human plus AI, not humans versus AI. AI is excellent at scale, repetition, and accelerating discovery, things like enumeration,

coverage expansion, and identifying patterns across large environments. Humans are still better at creativity, adversarial thinking, chaining complex attack paths, and understanding nuance and business workflows. In AppSec specifically, some of the highest impact findings still come from human intuition and experience. AI helps our team spend less time on repetitive work and more time on high value analysis. Thanks. And just to kinda call out a relevant start.

So you mentioned about, coverage. So I just wanted to call out something, related to that. So based on a twenty twenty six Omnia research study, we found that most enterprises are only testing about thirty two percent of their attack surface. So that leaves a really large sixty eight percent of critical gaps exposed. I think particularly when you think about, offensive AI can rapidly map an attack surface, identify weak points, and iterate on that at machine speed, everything becomes a potential attack vector.

Jessica, really interested in your your thoughts here. Why do you think that is that that there is such a large gap of sixty eight percent, and what does that mean in terms of business risk? From a coverage perspective, right, before testing before testing was often driven on annual pad tests, prerelease assessments, or even compliance focused exercise. The challenge is that the environment changes long before the next test cycle begins.

After moving toward continuous testing, the mindset the mindset changes completely. Security becomes part of operational rhythm instead of a checkpoint. Team starts prioritizing exposure management, real time validation, and ongoing attack surface awareness. Operationally, that means tighter integration between AppSec, engineering, cloud, SOC teams, along with tooling that supports continuous visibility instead of onetime assessments.

The turning point was realizing the scale problem wasn't going to go away. The attack surface kept growing faster than traditional testing models could keep could keep up with. What stood out with AI plus human approach was that the ability to expand coverage and accelerate testing without sacrificing validation quality. We saw that AI could help identify and prioritize opportunities much faster while human testers still provide the depth and context needed for meaningful findings.

Thanks, Jessica. It's a really important point that you mentioned about the both the sort of continuous testing side in terms of, like, that threat exposure time, which you need need to minimize as as you were mentioning, but also in terms of understanding what assets might be out there that you're not testing. So one of the things that we do at SYNNAC, and it's just shown in the top right on this this slide here, is that we provide attack surface discovery data in the portal, enabling unknown assets to be uncovered, and then we can rapidly use that information to start the

testing rapidly so we can sort of then minimize that threat exposure time. So just a couple of questions then related to that. What does expanding coverage look like in practice to you, and how does that change your team's workflow, tooling, or mindset when you make that shift from periodic to continuous testing? What's that before and after look like? Operationally, you gain much better visibility into assets and exposures that weren't per previously part of the conversation.

That changes prioritization, ownership, and remediation workflows pretty quickly. One of the biggest surprises is usually not sophisticated zero days. It's how many risks come from forgotten assets, exposed services, misconfigurations, or applications or applications that evolve faster than governance processes. Expanding coverage often reveals gaps between what organizations think they have and what's actually reachable

from an from an attacker's perspective. Thanks, Jessica. And just, like, question related to that then. When you expanded your attack surface, the coverage of that anyway, what what changed operationally? What what were some of the things that really surprised you the most in those findings when you when you covered much more of that attack surface? Were there certain things that's kinda stood out for you? It helped us definitely identify, other areas within our products and platforms that

we didn't see from traditional tests. And I think more visibility in that kind of area is what's going to help us not only secure our products, but also help us work with our teams even closer. That's a really good point. Thanks. And I think that really sets up what I wanna show next as well. So you've described what this looked like from the outside for you and that sort of path you went through,

but I just wanna give everyone a look at what Sarah's actually doing and how that works as well. So I just got a couple of slides to show that now. So Sarah, which is our Synak Autonomous red agent for pen testing. So this is a finely tuned agent harness which is built to attack and exploit attack services at scale safely and with quality. And that safely bit is really, really important as well. So just kinda talk about about how this works across the five

stages, which this breaks up into. Firstly, there's reconnaissance. So we undertake discovery based on the scope that's provided. Recon agents then decompose this into targeted tasks across web, database, network recon, depending on what was found. Those recon tasks then run-in parallel, and each of those is investigated independently. And this is so that we can reduce the overall assessment time. The information found can then inform the next stage for then how those attacks are gonna then proceed.

So we've then got a number of, attack vector specialists. So these are specific agents that can correlate weaknesses using TTPs and CVE intelligence, and these attacks are based on the findings from the previous stage and look at what attack paths are likely. These specialist agents then perform that deep testing, each focused on a particular protocol or technology, and this really aligns with that concurrent and robust testing. Then the third stage, we move on to the verification agents.

So we've then got a number of specialist verification agents each paired with the attack agents. These independently retest each of those findings. And this is a really important step because we want to ensure that they are actually exploitable when they find an exploitability and really eliminate those false positives. So this ensures both accuracy and reliability of the assessment. And then once we've got the information, we then move on to the reporting and strategic insight stage. And this is where report findings are synthesized with

all of the various information, whether that's steps to reproduce, whether that's the context. And each of those vulnerabilities of providers, each individually actionable. So that means people can see each exploitable vulnerability in its own right with associated patch verification. So that essentially means as you're fixing, as you're remediating, you can go in and you can retest that vulnerability. And a really important point to add here as well that each of those vulnerabilities that's been provided as part of the

reporting stage has had human validation. So this is making sure that they're all meeting a high quality bar and are having zero false positives. And then there's a full consolidated pen test report write up, making sure that this is hitting the same level of qualities you would expect to get from a human pen test. And this is what mirrors that same typical flow of what you get from a traditional pen test report. So just to show, as a diagram,

to show how this this flow is orchestrated. So each of these phases flows into each other. So you got reconnaissance flows into the attack agents, which flows into verification and then into reporting. But I just wanna just call out a couple of additional points to add as well. Firstly, this is all overseen by a central orchestrator, and this oversees the entire Sarah pen test end to end. So this maintains the global state. It coordinates each of those specialized agents, and it manages the overall strategy and resource allocation.

I'm working alongside this of rigorous guardrails that ensure that all of the testing is done with really safe behavior. Give you an example. We got content filtering, which would prevent a dangerous command such as drop tables. And when the findings come through, as I mentioned, all of those reports then go via vulnerability operations team, and we apply exactly the same quality checks to the reports from the AI agents from Sarah as what we do from the Synap

Red team. So just to kinda close that out, so Sarah's been built for the enterprise. So we've got audit trails, compliance, actual reporting. There's analytics, and then there's full integrations in RBAC all consolidated within the same platform. So just to show you a couple of screenshots from the Synap platform just to kind of illustrate that flow. So, firstly, the testing can be started fully on demand, so you can provide a scope manually.

But, additionally, talking with Jessica a minute ago about kind of understanding the attack surface coverage and how much has been tested, You can use the attack surface open and source intelligence data to drive that testing, to initiate the testing from. So you can use that to kind of accelerate how quickly you onboard those assets. And then high high quality reports are then provided back within a matter of days, all human readable, and, as we said, all quality checked as well.

And not shown here as well, but there's also a number of business intelligence analytics which run over top of this. So we talked about as you might be increasing the attack surface coverage. So you're gonna expect there's a lot more data that you've gotta look through. So we provide a number of analytics which can really help with that sort of root cause identification and see some of the patterns from that testing as well. So just to call out one anonymized case study for a recent test that we did.

So this is an example for some of the findings. I'll just I'll drill into some of the detail in a moment, but I just wanted to call out a couple of stats which I thought was really interesting. Firstly, the speed. So it's able to uncover these critical vulnerabilities in just six hours. And then the criticality of these exploitable vulnerabilities having gone through the full cycle from reconnaissance, the attack agents, the verification,

just really short time. Just yeah. I I find that quite fascinating how quickly that is. And the criticality of those findings, seventy percent critical or high. So then, yeah, let's let's just talk through a couple of, of of few of those examples. So three major vulnerabilities were found, and these had previously been pen tested as well. So that's quite shocking. They'd previously undergone a pen test from someone else before, and three major vulnerabilities were found. So Sarah was shown to be reasoning like an attacker and

mimicking human like offensive security thinking during this testing. So first first vulnerability on here, account takeover. So Sarah was able to, through looking at the forgotten password link, identify that there was a mistake in the reset token in the body of the information that was being returned, and this was visible to anyone. And through this, Sarah was able to confirm it could then log in with new credentials that it was able to change. Really impressive. It was not able just to identify it.

It could verify and prove that with impact as well. The second one, a SQL injection vulnerability. This is found not on the user interface of the application, but on the client side script. Initially, could find from reconnaissance on a sort parameter. And this enabled exposing every record in the system from credentials, roles, email addresses, everything needed to map out and compromise the user base. And then thirdly, a stored cross site scripting.

So this is where a persistent payload could be planted that could execute in every victim's browser, which could silently steal tokens and then a enable further takeovers. So each one of these vulnerabilities was really, really serious on its own. But combined, they represented a complete organizational compromise. You could enumerate all users, take over their accounts, and plant plant persistent code. This is just one example. But for me, I think just seeing the sort speed at which this could

operate, I thought was, yeah, quite impressive, certainly from my view. Jessica, I don't if you had any thoughts from anything I've gone through here just before, yeah, move on to the next one. No. I think you you hit the nail on everything, James. I think the perspective on the SYNNEX, SARA AI pen test, the initiative itself is really, it's really well put together. And I think the more that we continue on to use this type of service with us,

we'll definitely see we'll definitely have an orange to orange kind of comparison of what we've done in the past with our pen test. Definitely. So just let's talk a little bit about your mean time to remediation there, Jessica. So with increased scale in exploitable vulnerabilities, how does your team prioritize and decide what to act on first, particularly when vulnerabilities are identified at scale?

One of the biggest lessons for us was realizing that scale alone does not solve security problems. Prioritization does. Back in twenty twenty one, we had visibility into growing number of findings, but remediation time lines were still too long because teams were overwhelmed with volume and inconsistent risk context. What made the biggest difference was shifting from vulnerability counts to exploitability and business impact.

We focus heavily on validating what was actually actionable, externally exposed, or realistically expose exploitable. That help engineering teams spend time on the issues that matter most instead of chasing noise. Operationally, we also improve collaboration between abstract infrastructure and development teams, standardized workflows, and created tighter feedback loops around remediation tracking.

Over time, that reduced friction and significantly improved response times across the organization. Thanks. That's really, really interesting. So I know, yeah, you were able to reduce your mean time to remediate by ninety eight percent, and I'm sure that took a lot of effort on your side to sort of build out and coordinate that across your organization. And as you're saying, you started those sort of efforts back in twenty twenty one, and back then, your remediation time was three

times what it is today. If you were just to kinda call out a couple of main actions or just one or two main things that you think that you did that made that sort of biggest difference for you, what do you think those main points would be? Yeah. I think, like I mentioned, the first one will be a lot of collaboration with with, teams across the board. Right? With that, building a channel a communication channel with everyone, making sure

we are up to date, we are all on the same page. And lastly, set up a cadence. I know in the beginning, it could be a little bit of a time consumption, but setting up a cadence to understand workflows from not just my end, but also from our product teams and even, any other, dev teams, network teams, just to make sure that we are all aligned.

We we keep those cadences. And once we have kind of a a structured workflow, we're like, okay. Fine. We'll we'll we'll pull back on the cadences. We have a we have a routine. We have a rhythm, and let's just go with what we built. So I think the first thing is just always collaborating. Right? I think, information security in the beginning when I first joined this world, I was always told,

things are siloed, and I didn't like that. I'm like, no. We gotta we gotta work together. You know, security is everyone's responsibility. And I feel like if we hold off in not sharing and communicating what's really going on, then those are the gaps that won't be filled. Yeah. Communication is definitely really important, isn't it? But, yeah, I think kudos to yourself. I think that ninety eight percent reduction is yeah. That's really impressive. Thank you. So, yeah, a lot of,

security teams worry that more coverage just means more noise, particularly when considering that it might be inundated by false positives. From your perspective, how does the human validation of or triage component change that equation for you? And what would your what would your program look like if you didn't have that human validation? So that that concern is valid because expanding coverage without validation can absolutely create alert fatigue.

More findings don't automatically mean more security value. Human validation changes the equation because it adds context, credibility, and prioritization. It helps separate theoretical issues from practical risk, gives engineering teams confidence that what they're fixing actually matters. Without that human element, the program will likely become very tool heavy and volume driven,

which can lead to teams disengaging over time. The combination of AI driven skill and human validation gives us both coverage and trust, and that trust is what makes remediation move faster. Yeah. Trust is definitely really important, and I I expect I'm interested in your thoughts here. When you were kind of driving that MTTR time down, I'd imagine trust was a key part of that.

Is that right? Correct. And just finally, what's some key steps that you think teams can take to start testing more of their attack surface overall? The first step is gaining accurate visibility into what actually exists in your environment. You can't secure what you don't know about. Start by identifying Internet facing assets, APIs, cloud services, and business critical applications.

Then prioritize continuous testing around exposure and risk rather than waiting for annual cycles. The goal doesn't have to be perfection on day one. It's building a process where coverage continuously improves over time. Thank you, Jessica. Really, really insightful. Just wanna say really appreciate hearing from your experience, insights. I've certainly found it really, really intriguing.

If if what Jessica's shared here today resonates with you and you're ready to see what your attack surface and risk actually looks like, please head to Synap dot com to sign up for your free SARA pen test. You might be surprised what's been sitting there untested and the level of risk exposure that those assets present. Thank you all for coming, and thanks again, Jessica. Thank you, James. Take care.

Speakers

Jessica Boy

Paramount

Lead Information Security Analyst

Solutions Architect, UK & Ireland

AI pentesting and attack surface coverage FAQ

What does good security coverage mean now that AI has changed the attack?
Jessica Boy's answer in the session: AI moved the goalposts on what counts as good enough. Testing used to be periodic and aimed at known critical assets. Applications, APIs, cloud services and AI integrations now change faster than snapshot-based testing can follow, so coverage means continuous visibility across a much larger attack surface, including shadow APIs and AI-enabled features. The question is no longer whether something was tested, it is how quickly risk can be identified and validated as the environment changes.
Why do most enterprises test only part of their attack surface?
Research conducted with Omdia found that the average enterprise tests about 32 percent of its attack surface, leaving 68 percent untested. The reason given in the session is structural rather than negligent: testing has been driven by annual pentests, prerelease assessments and compliance exercises, and the environment changes long before the next test cycle begins. Expanding coverage usually reveals the gap between what an organization believes it has and what is actually reachable from an attacker's point of view.
Can frontier AI models replace human penetration testers?
Both speakers say no, for the same reason. Frontier models are force multipliers that improve scale, coverage and speed in reconnaissance, pattern analysis and identifying potential attack paths. Human expertise stays critical for validating business logic issues, understanding real risk impact and separating noise from meaningful findings. As Jessica Boy puts it, the strongest model is human plus AI, not humans versus AI, and in application security some of the highest impact findings still come from human intuition and experience.
How does Sara AI Pentesting actually work?
James Duggan walks through five stages. Reconnaissance agents decompose the provided scope into targeted web, database and network tasks that run in parallel. Attack vector specialists correlate weaknesses using TTPs and CVE intelligence, each focused on a particular protocol or technology. Verification agents, paired with the attack agents, independently retest each finding to confirm it is genuinely exploitable. Reporting synthesizes each vulnerability with steps to reproduce, context and patch verification so it can be retested as it is fixed. A central orchestrator maintains global state and resource allocation across all of it, with runtime guardrails such as content filtering that blocks destructive commands.
Does expanding coverage just create more noise for the security team?
That concern is valid, and the session addresses it directly. Expanding coverage without validation creates alert fatigue, and more findings do not automatically mean more security value. Human validation adds context, credibility and prioritization: it separates theoretical issues from practical risk and gives engineering teams confidence that what they are fixing actually matters. Without it, a program tends to become tool heavy and volume driven, and teams disengage over time.
Where should a team start if it wants to test more of its attack surface?
Start with accurate visibility into what actually exists: internet-facing assets, APIs, cloud services and business critical applications. Then prioritize continuous testing around exposure and risk rather than waiting for annual cycles. The goal is not perfection on day one, it is building a process where coverage keeps improving over time.

Next step

Find out what is sitting untested on your attack surface.

Run a Sara AI pentest against a defined scope and see which findings the Synack Red Team confirms as exploitable. Most teams are surprised less by the sophistication of what turns up than by how long it has been there.