AI Pentesting Works. Building It Yourself Is the Hard Part.
AI pentesting works, but building it in-house is the hard part. Synack VP Chris Brown breaks down the reliability, token economics, model dependency and validation costs that come with building versus buying an AI pentesting capability.
Key Takeaways
- Building AI pentesting in-house means owning every cost that comes after the demo, including reliability engineering, human oversight and validation infrastructure.
- A one-time proof of concept does not reveal real economics. Tokens, compute, orchestration and monitoring costs scale with every additional test.
- A specialist vendor spreads engineering, orchestration and validation costs across an entire customer base; an in-house build carries them alone.
- Building on a third-party foundation model means inheriting that vendor's pricing changes, deprecations and lifecycle decisions as a permanent dependency.
- The model is not the product. Orchestration, safeguards, testing infrastructure and human validation are what make an AI pentesting result trustworthy.
- The build-vs-buy decision should weigh specialist talent, infrastructure, maintenance and validation costs, not just the initial development price tag.
AI pentesting is already changing security testing. It can increase speed, coverage and testing frequency in ways that simply were not practical before. The opportunity is real, and I think the change is permanent.
The mistake I see companies making is different. They assume that because they can access the same foundation AI models as a security vendor, they can therefore build the same capability themselves.
That’s where the economics start to get interesting. I’ve spent more than 20 years in cybersecurity and I’ve seen plenty of technology cycles come and go. Some genuinely reshape the market. Others create a lot of noise before commercial reality catches up. AI may well be one of the biggest shifts we have seen. But that does not mean we should suspend commercial judgment about what it actually takes to operate it.
The AI Demo Is Not the AI Business Model
One of the easiest mistakes to make with AI is being impressed by the first result. You give an agent a task and it produces something in minutes that would have taken a person hours. It looks good. It feels transformative.
Then you try to operationalize it. Can it do the same thing tomorrow? Can it do it next week? Can you trust the output consistently?
What happens when the underlying model changes? What happens as the workflows become more complicated? What does it cost when you start running it hundreds or thousands of times rather than once?
Those are very different questions.
I have seen AI tools work brilliantly at the beginning and then start to drift. The output changes. Instructions are interpreted differently. Something that looked reliable as an experiment becomes much harder to trust as an operating process.
That matters in any business. It matters even more in cybersecurity. And if you have built the capability yourself, that gap between demo and operating process is now yours to close.
Building Reliability Is More Expensive Than Building the Demo
People tend to discuss AI reliability as a technical problem. I see it as a business problem.
Getting an AI agent to produce an impressive first result is one thing. Building a security capability that produces dependable results repeatedly across thousands of tests is another.
If you build internally, your team owns that problem.
If a human has to keep checking whether the agent has done what it was supposed to do, that oversight has a cost. If results need validation, that has a cost. If you need additional infrastructure, orchestration and engineering to make an AI workflow dependable, that has a cost. And if the people you serve do not trust the output, you have a much bigger problem.
This is not only a vendor’s view. The NIST AI Risk Management Framework frames AI trustworthiness as something organizations need to build in across the design, development, use and evaluation of AI systems, not treat as an afterthought once something is in production.
Automation only creates value when the result is useful. In cybersecurity, speed without confidence is not enough.
Your Proof of Concept Does Not Show You the Real Cost
There is another assumption I hear constantly: AI is going to make everything cheaper. Maybe. But show me where that data comes from.
The economics of large-scale AI are still developing. Tokens cost money. Compute costs money. Orchestration, testing infrastructure, monitoring, engineering and validation all cost money. Repeated agentic workflows cost money. More sophisticated testing requires more compute, not less. A proof of concept run once tells you almost nothing about that.
In The Hidden Costs of Building an AI Pentesting Solution, we’ve mapped exactly where those costs accumulate, from agent tokens and compute to engineering effort, model changes and ongoing maintenance. I will not repeat that analysis here, but my point is the commercial one that sits on top of it.
The price a customer sees today does not necessarily tell you what the long-term economics look like. A company can price aggressively to acquire customers. A well-funded vendor can subsidize growth. A platform provider can make usage extremely attractive while the market is developing. That is a perfectly legitimate business strategy. It is not the same thing as proving that the underlying economics work at scale.
There is a difference between a cheap initial test and a sustainable enterprise capability. If a team loves running one AI test for a small amount of money but the economics do not hold when they need 100 of them, that is worth knowing before you commit to building.
A specialist vendor spreads those investments across a platform and a customer base. An enterprise building internally carries them itself.
If You Build It, You Own the Model Dependency Too
This is the part I think many companies are underestimating. Most enterprises are not going to train their own foundation model for pentesting. They will build on somebody else’s. That means your internal solution inherits another dependency: model availability, API changes, pricing, deprecations and performance changes. Your workflows, integrations, prompts and engineering start to depend on someone else’s platform decisions.
This is not hypothetical. AI providers retire models on their own timelines. Anthropic, for example, documents a formal model lifecycle of active, legacy, deprecated and retired states, and notes that applications may need to be updated as models are retired.
A vendor building AI security technology for a living has teams whose job is to manage that complexity: testing new models, absorbing pricing changes, re-tuning when behaviour shifts, keeping the service dependable through all of it. If you build it yourself, it becomes your team’s job.
Technically, you may be able to change providers. Commercially and operationally, that is much harder once your product has been built around one ecosystem.
What Are Security Buyers Actually Trying to Achieve?
None of this changes the fact that customers are asking for AI. That signal matters. If a customer comes to us asking for AI pentesting, I am not going to start by telling them they are asking the wrong question. I want to understand what they are actually trying to achieve.
Do they want more coverage? Faster testing? Lower cost? More frequent testing? Less reliance on scarce skills? A different operating model?
That is where the real conversation starts. The customer may use terminology we would not use internally. Fine. The objective is not to win a terminology debate. The objective is to understand the pain behind the request. AI has absolutely changed that conversation. What it has not done is remove the need to ask whether the solution is reliable, whether the customer will trust it and whether the economics stack up, whether you buy it or build it.
Why Fidelity Matters in AI Pentesting
Right now, a lot of the AI security conversation is about speed and automation. I think the conversation will increasingly move toward fidelity.
Did the testing find something meaningful? Can I trust the result? Can I distinguish noise from real exposure? Can somebody prove that a vulnerability is exploitable? Can I make a business or security decision based on what I have been given?
The question is not whether AI belongs in pentesting. It does. The question is what you have to build around the AI to make the result useful: orchestration, safeguards, validation, testing infrastructure and, where it matters, human expertise. That complete system is the product. The underlying model is only one component.
The model is not the product.
How Should Security Leaders Think About Build vs. Buy?
If you’re deciding whether to build AI pentesting internally or buy an established capability, the initial development cost is rarely the real cost. Everything above, reliability, token economics, infrastructure, model dependency and validation, becomes part of the total.
Synack explored exactly that question with Dow’s cyber engineering lead Dan Lacher and Synack CTO Mark Kuhr in the Build or Buy? The AI Pentesting Question Every Security Team Is Facingwebinar. Take a look to learn why Dow concluded that token consumption and engineering effort made an internal, production-grade agent less attractive than partnering.
Separate the Market Shift From the Hype
AI is not going away. Buyer expectations have already changed. Cybersecurity vendors that ignore that will have a problem.
But there is a difference between recognizing a real market shift and assuming every organization should build its own capability around it.
I have seen this cycle before. New technology arrives. Funding follows. Expectations accelerate. Prices get aggressive. Everyone wants to claim the new category. Then the market starts asking harder questions. Does it work consistently? Do the results hold up? Can it be trusted? Can it be delivered profitably?
Those are the questions that ultimately separate durable capabilities from hype. AI is changing pentesting, and I think that change is permanent. But access to an AI model does not mean every enterprise needs to become an AI security software company.
Before you decide to build, ask a harder question than “Can we create this?” Ask whether you want to own the reliability, infrastructure, model dependency, validation, maintenance and economics that come with it. Because building the demo may be surprisingly easy. Building the business-grade capability is the hard part.
Check out Sara AI Pentest to see how Synack combines AI-led coverage with human validation, so you do not have to build all of that yourself.
Related reading: Build vs. Buy AI Pentesting: Why Dow Chose to Partner With Synack • Continuous Security Validation: Why Synack Built for It
Watch the Build vs Buy webinar
Dow's Cybersecurity Engineering Team evaluated several AI pentesting tools and considered building the capability internally, but they ultimately chose Synack.
Frequently Asked Questions
Agentic AI workloads can involve model usage, compute, orchestration, testing infrastructure, monitoring and human oversight. The economics of a proof of concept can therefore look very different from running AI pentesting repeatedly across an enterprise environment, and an internal build carries all of those costs itself.
Model behaviour and application performance can change as models, prompts, integrations or underlying workflows change. Security teams therefore need ways to monitor repeatability, reliability and the quality of results rather than assuming that an initial result will remain consistent.
Yes. Products built deeply around a particular model or API may require engineering work when providers update, deprecate or retire models. Anthropic, for example, formally maintains active, legacy, deprecated and retired model lifecycle stages. If you build in-house, managing that dependency becomes your team’s responsibility.
The decision should include more than the initial development cost. Teams should consider specialist talent, model and infrastructure costs, maintenance, validation, compliance requirements and whether building the capability is strategically core to the business. Synack’s build vs. buy guidance goes deeper into these considerations.
AI can expand the speed and breadth of testing, but security teams ultimately need confidence that a finding represents real, exploitable risk. Synack’s AI pentesting approach combines AI-driven coverage with human validation for that reason.


