Offensive Security Is Becoming a Program, Not a Purchase
You know what you spent on penetration testing last year. But can you explain what that investment covered between engagements? For many organizations the honest answer is that it covered the weeks the testers were working against a scope agreed before they started. That was a reasonable arrangement when systems changed a few times a year. It is worth revisiting, now that many of them change weekly.
Key Takeaways
- A penetration test buys an assessment of an agreed scope over a defined period. That remains valuable, and it is not the same as coverage between engagements.
- Additional engagements can improve coverage. They do not automatically improve coordination or remediation, which is where value is often lost.
- Moving to a program changes scope governance, funding, internal accountability and how suppliers are evaluated.
- Independent research points to the same conclusion: technology alone does not resolve remediation bottlenecks. Workflow change and shared accountability are required.
- Testing obligations vary by framework, entity and scope. A continuous program has to be designed against the requirements that apply, and it does not automatically replace prescribed assessments.
What Is an Offensive Security Program?
An offensive security program is an ongoing operating model for finding and validating exploitable risk, with defined scope governance, continuous funding, named accountability for remediation, and evidence that carries across periods. A penetration test is one engagement inside that model. The program is what determines whether coverage holds between engagements.
When a security team commissions a penetration test, the line item says testing. What the organization receives is an assessment of an agreed scope, carried out over a defined period, and a report.
That is a valuable thing to buy. Specialist engagements, red team exercises and point-in-time assessments all have a place, and in several contexts they are specifically required. The question is not whether to keep them. It is whether they are the whole of what the organization needs.
It is worth checking how your own arrangement actually works. Is the scope fixed well before the work begins? Is the budget approved annually? Is the provider assessed mainly on whether the engagement ran to schedule and whether the report was thorough? Those can be reasonable measures of an engagement. They say very little about the periods between engagements.
Why More Engagements May Not Close the Gap
When coverage feels thin, the common response is to buy more testing. A second engagement, a wider scope, an additional supplier for the areas the first does not reach.
That genuinely improves coverage. What it does not automatically improve is coordination and remediation. Findings arrive from more sources, on different cycles and in different formats, and someone still has to decide what gets fixed first and who does the fixing.
More findings are not necessarily more progress. A scalable offensive security program
must distinguish between possible exposure and proven exploitability, so security teams
can focus remediation on the risks most likely to create real-world impact. AI can
expand testing frequency and coverage, while human expertise provides the context,
business-logic analysis and validation needed to prove what matters.
Accountability between engagements is more often unclear or fragmented than absent. Application teams own their code. The security team owns the testing relationship. The question of what shipped last month, and whether anyone assessed it, tends to sit in the space between them.
This is not only a supplier problem. A public G2 reviewer of Synack raised a useful limitation: keeping an assessment productive over successive years requires customers to update scope and context.
That describes recurring work on the customer side. It needs an owner, and naming that owner is a decision the buying organization has to make for itself.
What Changes
POINT-IN-TIME ENGAGEMENT VS. OFFENSIVE SECURITY PROGRAM
| Dimension | Point-in-time engagement | Offensive security program |
| Scope | Fixed and agreed before work begins, typically reviewed once a year. | Updated on a defined process, with named approvers and a record of what was in scope when. |
| Cadence | A defined window, scheduled in advance. | Testing recurs with change, release and risk, not the calendar alone. |
| Funding | Approved per engagement, usually an annual line item. | Funds an ongoing capability, including remediation and retesting. |
| Accountability | Report is delivered to the security team. Ownership after delivery is often unstated. | A named owner for findings that arrive mid-release, agreed in advance with application owners and release managers. |
| Evidence | A report covering one scope and one period. | A record of scope, severity and closure that holds across periods and regions. |
| Remediation and retest | Often out of scope or purchased separately. | Part of the operating model, with closure tracked through retest. |
| Supplier evaluation | Tester expertise, schedule adherence, report quality. | Adds workflow integration, coverage across relevant surfaces, and follow-through on unfixed issues. |
| What it proves | The assessed scope was tested during that window. | Whether exploitable risk is being found, validated and closed as the environment changes. |
Four areas change when offensive security moves from a series of engagements to an ongoing program, the shift now described as continuous offensive security. Testing methods and delivery models can change as well. These four are the ones that determine whether the investment works.
Scope
Scope stays documented, authorized and controlled. What changes is how often it is reviewed and how quickly a change in the estate reaches it. Rather than a scope written once a year, you need an agreed process for updating it, a named approver, and a record of what was in scope and when. That is a governance and contracting question as much as a security one.
Funding
Funding needs to cover an ongoing capability, including remediation and retesting, rather than a single assessment. That is a different budget conversation and not necessarily an easier one. Ongoing investment has to be justified by evidence of value, which means deciding at the outset what you will measure and what would count as progress.
Accountability
Someone has to own what happens when a finding arrives in the middle of a release cycle. No supplier can answer that for you. In practice it means the security team, the application owners and whoever controls the release process agreeing in advance how findings are triaged, prioritized and tracked to closure.
Supplier evaluation
Tester expertise remains a valid and important criterion, and it should stay on the list. Alongside it, ask how findings integrate with your development and ticketing workflows, what coverage the program provides across the surface you care about, and what follow-through looks like when something is not fixed.
What the Research Says, and What I Take From It
THE COVERAGE GAP IS ALREADY VISIBLE
Synack’s 2026 State of Continuous Security Validation found that 95% of enterprises
discovered high or critical vulnerabilities outside scheduled testing windows during the
previous year. Thirty-eight percent said at least one-quarter of their critical attack
surface had not been independently tested or validated within the previous 90 days,
while only 15% described their security validation program as continuous.
Independent analysis points in the same direction. Gartner’s Hype Cycle for Security Operations, 2026 describes a shift toward proactive, continuous validation. It also makes a practical point: technology alone cannot resolve remediation bottlenecks. Organizations need to change workflows and establish shared accountability with the teams responsible for fixes.
Gartner’s guidance on penetration testing as a service places similar emphasis on appropriate testing scope, hybrid models combining human expertise with automation, alignment with compliance obligations, and integration with development and ticketing workflows.
For me, the commercial implication is clear: buyers need to evaluate how testing fits into their organization, alongside the quality of the testing itself. A provider can deliver more testing. The organization still has to decide who acts on the findings and how that work is funded. Those responsibilities are better agreed before a contract is signed than discovered afterwards, because a recurring subscription can leave the same gaps as an annual engagement if nothing else changes.
The trade-off is worth stating plainly. More testing produces limited value if the organization cannot prioritise and remediate what it finds. Buying more of the input does not by itself improve the outcome.
A Market Signal Alongside the Research
On September 2, NetSPI and Synack announced a definitive agreement to merge,
subject to customary closing conditions and regulatory approvals. Synack CEO Jay
Kaplan explains why human expertise remains essential alongside AI in Autonomy Was
Never the Goal. Synack CTO Mark Kuhr explores how AI and expert testing should be
combined in Putting Experts on the Problems Only Experts Can Solve.
The proposed combination reflects the same market shift described in this article:
enterprises increasingly need testing depth, continuous coverage, integrated workflows
and evidence that persists across engagements. Once the transaction closes, the
strategic value of the combined organization will be measured not simply by the breadth
of services available, but by how effectively those capabilities can support customers as
one coordinated offensive security program.
Two Practical Considerations
Compliance
Testing obligations vary by framework, by entity and by scope. A continuous program must be designed to meet the requirements that apply to your organization. It does not automatically replace prescribed assessments or reports, and continuous testing does not by itself demonstrate compliance across a period. Evidence has to match the relevant controls, the defined scope and the required timing. That is a design decision to take with your compliance function early, not a property you acquire by changing testing cadence.
Operating across markets
If you run offensive security across regions, a single global program needs two things working together. Common governance, so that scope definitions, severity ratings and reporting mean the same thing in every market. And a current mapping of the obligations that apply in each of those markets.
The mapping can fall between legal, compliance and security rather than sitting clearly inside any of them. It is worth confirming who owns it, and whether it feeds directly into scope and evidence rather than living in a separate document reviewed once a year.
Questions to Answer Before the Next Investment
Five, and they are deliberately uncomfortable.
- What did we change or ship in the last twelve months, and how much of it was assessed?
- Who owns exposure between engagements, by name rather than by function?
- What happens to a finding that arrives in the middle of a release?
- What evidence do we hold, for which scope and which period?
- If this works, what do we stop funding, and what might we need to fund instead?
The last question deserves an honest answer in both directions. Expanding a program can mean spending more rather than less, and that case has to be made on evidence like any other.
How Synack Approaches This
Synack supports offensive security as an ongoing program rather than a disconnected
series of engagements. The Synack Red Team brings human expertise to complex
attack paths, business logic and real-world adversarial testing, while Sara AI Pentesting
expands the frequency and coverage of testing.
Findings are validated for exploitability before they reach the customer, helping security
teams focus on proven risk rather than an unprioritized queue of possible exposures.
Scope, testing activity, validated findings, remediation status and retest evidence are
maintained within the Synack Platform, creating a record that remains usable across
reporting periods, business units and regions.
This combination of AI-powered scale and human validation enables continuous
pentesting at scale: AI finds more. Humans prove what matters.
The Decision Worth Making
Before approving the next offensive security investment, agree what it covers, who owns remediation, what resources they need and how progress will be measured. Those decisions turn testing into a program. They do not require you to change supplier, but they do require ownership.
For details of the announced agreement, read NetSPI and Synack Are Coming Together. To discuss what an ongoing program looks like in practice, talk with our team.
Frequently Asked Questions
At minimum: a maintained scope with named approvers, testing that recurs with change rather than only on the calendar, validation that confirms which findings are actually exploitable, named ownership for remediation, retesting to confirm closure, and evidence that holds across periods and regions. The governance around the testing matters more than the tooling.
Posture changes with every release, so the gap between engagements is where most unmeasured risk sits. Closing it takes testing that recurs with change, validation of which findings are exploitable in the current environment, and a closure record showing what was fixed and retested. Coverage between tests is a governance question before it is a purchasing one.
A penetration test buys an assessment of an agreed scope over a defined period. A program buys an ongoing capability: scope that is maintained, testing that recurs, remediation that is owned by name, and evidence that accumulates. Buying more engagements increases coverage. It does not by itself improve coordination or closure.
Not automatically. Testing obligations vary by framework, entity and scope. A continuous program can produce stronger evidence than a single annual assessment, but it has to be designed against the requirements that apply, and it does not by itself replace a prescribed assessment or demonstrate compliance across a period.
Agree it before it happens. The workable model is a named owner for triage, a severity and prioritization standard shared with application owners and release managers, and a closure path that includes retest. Ownership defined at function level, the security team, tends to leave mid-cycle findings sitting between teams.
Define the measures before the investment: how much of what changed in the period was assessed, how many findings were validated as exploitable, time from finding to retested closure, and how much of the attack surface the current scope covers. Progress is the movement in those numbers, not the number of reports received.


