Putting Experts on the Problems Only Experts Can Solve

Expert judgment is the scarcest resource in offensive security. Making it accessible to more of the world, continuously, is an engineering problem.

Key Takeaways

  • More testers and more agents don't automatically mean better security outcomes.
  • Routing only works when context and evidence travel with the work, and the escalation rules are explicit.
  • A model that crushes a benchmark can still get lost in a live production environment.
  • Validation can't wait until the end. It has to happen at every handoff.

Jay and I started Synack after working together at the NSA with a straightforward goal. The people who can think like attackers are scarce, and most organizations get almost none of their time. We wanted to make that expertise reachable for far more of the world, and available continuously rather than once a year in a two week window.

That is still the problem. It is the reason AI matters here at all. Not as a substitute for expert testing, but as the only way to put expert-led testing in front of the number of organizations and the amount of attack surface that actually need it.

Today, NetSPI and Synack announced a definitive agreement to merge. Jay laid out why the combination is a bet on expertise at AI scale. I keep landing on a narrower question. As coverage expands, how do we make sure expert time lands on the problems that genuinely need it, and that nothing is lost when work moves between people and the systems supporting them?

More Experts Changes What Is Possible

The easy read on any security merger is that a bigger bench means more testing capability. That is true, and it matters more than it usually gets credit for. Skilled offensive testers are the scarcest thing in this industry, and there are about to be a great many more of them working toward the same goal.

The question I spend my time on is what surrounds them. Expertise turns into coverage only when the work reaches the right person while it still matters, and when what they learn survives the handoff to whoever picks it up next. Get that wrong and you add coordination cost rather than security outcomes, no matter how good the people are.

Agents do not solve that on their own either. Point enough of them at an environment and you will generate a great deal of activity, some of it useful, while the attack paths that would actually hurt you sit untouched.

Where Expert Time Should Go

Take an authenticated enterprise application with several user roles, approval steps and dependencies on systems it does not own.

Some of that work rewards patience more than judgment. Enumerating endpoints. Replaying a test across hundreds of inputs. Grinding through known weakness patterns. It has to be done, it is rarely what a good tester wants to spend a week on, and an agent does not get bored.

Then the application behaves differently after an approval, or a role change, or some sequence of individually valid actions nobody designed for. Now you are in state and business logic. Somebody has to form a hypothesis about how two systems interact, then be wrong three or four times before they are right.

That is the work you want your most experienced people on, and it is the work there is never enough time for. Everything underneath exists to get them there sooner and with more context in hand.

Which comes down to what travels with the work. The system has to retain what has already been learned, notice when an agent is looping, recognize when a result calls for a different specialization, and hand over enough evidence that the next tester is not starting from a blank page.

Not whether the agent is good, but whether the system around it knows when to stop asking the agent and bring in a person.

Lab Performance Is Not Production Performance

I recently argued that the agent harness matters more than the model. This is why.

Benchmarks are useful, but they simplify away much of the real work. The agent gets a clean target, stable access and a clear finish line. Production testing rarely looks like that. Sessions expire. Rate limits kick in. Authentication breaks. Workflows cross systems, and the application may change while the test is running. A model can look excellent in a capture-the-flag exercise and still get lost inside a real application.

That is why the harness carries more weight than the model. It has to hold state, keep the test inside scope, recover when something breaks and know when to stop. And it has to move work between specialized agents and people without throwing away context on the way. Model capability matters. What you get out of it depends entirely on the system you build around it.

Validation Is Where Scale Backs Up

Together, NetSPI and Synack bring more than 13 million hours of real-world offensive testing. What matters is not the total. It is that the hours are not concentrated in one method or one kind of environment. Mainframes and cloud. OT and modern web apps. Staff consultants who live inside a client’s environment for weeks, and a global researcher community that hits thousands of targets in parallel. Two sets of experts who learned the same lessons in very different places, and who are rarely wrong about the same things.

That breadth changes which problem is hard. At this scale, finding more things stops being the constraint. Every new layer of coverage produces more observations, more candidate attack paths and more decisions about what deserves another look. Discovery is not the scarce resource. Judgment about what is real is.

Which is why validation cannot sit at the end of the pipeline. Save it for the end and everything piles up before anyone challenges any of it. Build it in earlier and the questions get asked while the answers still cost something to ignore. Can the behavior be reproduced? Is the vulnerable component actually reachable in the environment as deployed? Does the evidence support the impact being claimed?

Ask those at the point the work moves and the experience on both sides compounds. Leave them to the end and all you have built is a longer queue.

The Engineering Problem Behind the Combination

NetSPI brings deep, consultant-led testing across the environments most enterprises spend the most on, web, application, host and network, with additional specialist depth in areas such as mainframe and OT/ICS. Synack covers much of the same ground with a different model: testing led by the Synack Red Team and, increasingly, supported by agents, with continuous testing and AI-enabled coverage as its distinctive contribution. The two overlap on most of the surface. The difference that matters is the delivery model, and that is exactly what makes the engineering between them worth getting right.

None of this is about doing less expert testing. The problem we set out to solve has not changed since we started the company. Expert offensive security is scarce, it is expensive, and most organizations get far less of it than their attack surface warrants. AI is how that expertise reaches more of the world, continuously, instead of a handful of applications once a year. Closing that gap takes more expert judgment in the loop, not less.

The implementation work comes after close. For now the point is simpler. Bringing experts together is one thing. Building the system that lets them work at full stretch is another.

Merging teams is the easy part. Merging judgment is not. That is the engineering problem worth solving.

For answers about the agreement, timing and what it means today, read NetSPI and Synack Are Coming Together.

Learn how the Synack Platform can secure your organization