If you’re evaluating an agentic pentesting solution right now, you’ve probably heard the same pitch more than once: point it at a target, and it discovers, validates, and exploits attack paths autonomously, the way a real attacker would.
That promise is worth taking seriously. It’s also worth pressure testing, and three questions do the heavy lifting.
- What can the assessment actually prove?
- When is that proof produced? And,
- How much of your environment does the proof cover?
Most evaluations stop at the first step. However, it’s at the second and third ones where validation programs are won or lost.
One note on where we stand: Picus builds and sells autonomous pentesting. That is exactly why we can be precise about where it ends, because the limits belong to the method, not to any vendor's implementation, and no roadmap can remove them.
The problem in four numbers
Four numbers from this year explain why the second and third questions now carry more weight than the first.
- Volume: 35,364 CVEs found in the first half of 2026, up 49.5% year over year.
- Prioritization: only 95 of roughly 39,600 CVEs published through August saw confirmed in-the-wild exploitation, so severity-led prioritization is effectively chasing the wrong list.
- Speed: mean time from disclosure to exploitation collapsed from 21.5 days in 2025 to 8 hours in 2026.
- Capacity: of the more than 26,000 vulnerabilities surfaced by AI-scale discovery, only 421 were patched upstream.
You can’t patch, schedule, or predict your way out. You validate, on evidence, within hours of the change that demanded it. Yet an annual pentest leaves as much as a 365-day blind window between a change and the next test, and weekly automated runs still leave a gap of up to seven days.
Against an eight-hour exploitation window, both lose.
Gartner has drawn the conclusion formally: its Continuous Offensive Security Testing (COST) model replaces the point-in-time assessment with trigger-driven, risk-tiered testing completed within risk-aligned timeframes, often minutes or hours, with a planning assumption that by 2028 over 60% of enterprise pentest programs will operate as continuous validation. The key question is no longer "have we tested?" It is "how quickly are we validating new exposures?"
Agentic pentesting is the most important answer to that question.
Two proofs agentic pentesting delivers
Every pentest exists to answer one question: are we exploitable? Agentic pentesting answers it with the two proofs that matter. It proves individual exposures are exploitable or not, confirmed by safely executing the exploit rather than inferring from a version banner. And it proves chained reach: where real exploits are chained from initial access through privilege escalation and lateral movement, with the confirmed path to a critical asset reported as evidence.
On the assets it reaches, nothing is more convincing, and because every fix can be revalidated with one more run, remediation becomes defensible rather than hopeful.
The operative phrase here is "on the assets it reaches." Proving exploitability is the first question answered; the phrase quietly concedes the other two: when the proof arrives, and how much of the estate it covers.
The speed gap – the “when”
Run an agentic pentest across a 250,000-endpoint estate and the full cycle takes weeks. That is a dramatic improvement over the quarter a human-led engagement needs, but it's still the wrong unit of time.
Careful agentic pentesting works step by step: foothold, enumerate, pivot, chain, prove. Repeat that across a quarter of a million endpoints, and the sweep stretches out while the environment changes underneath it. By the time it completes, large swaths of the results describe an environment that no longer exists.
Weeks may be fast for a pentest, but against an eight-hour window, it's slow. A deep sweep alone can't bridge that gap.
The coverage gap – the “how much”
The coverage gap matters more. Live exploitation can only be pointed at part of the environment. Business-critical production systems, very large segments, and restricted or air-gapped zones are off-limits to real exploits for safety, stability and access reasons. The thousands of CVEs with no working exploit give the tool nothing to fire. And the window between disclosure and the first exploit, where the speed numbers show the entire race is now taking place, is precisely, and unfortunately, when an exploit-dependent tool has the least to say.
Add it up and autonomous pentesting alone sees perhaps 20 to 30% of real exploitability in a typical enterprise. Stacking a second or third tool doesn’t change the math, because every tool in the category shares the same method, and therefore the same ceiling. And partial coverage doesn’t reduce uncertainty. It simply relocates it, to the assets you skipped, and the attacker only needs the one thing you skipped for things to go pear-shaped.
The change selects the validation method
The coverage gap closes the same way: let the change that triggered the test select the method that answers it fastest and safest.
An emerging vulnerability lands on your assets. The key question is "is it exploitable in our environment?" Exploitability validation answers this within hours, across the whole affected scope, by testing the attacker techniques the vulnerability depends on, no working exploit required. This answers the objection we hear most often, "doesn't pentesting already cover that?" Yes, but only where live execution can go; in contrast, exploitability validation reaches most of the estate and most of the CVE list.
A new campaign is observed, or a security control changes. The question here is "can we stop this?" Security control validation emulates the campaign's techniques against your live defense stack and shows whether each is prevented, detected or missed.
An infrastructure change lands. The million-dollar question is "did this open a new attack path?" Agentic pentesting confirms it with chained, live proof, the thing only it can do.
Three methods, one findings model
One warning from recent history: vulnerability management splintered into islands of duplicate findings and conflicting priorities, and validation will recreate those islands if each method ships with its own console and queue. The three must feed one shared findings model, deduplicated, evidence-backed, and asset-aware, with one backlog and one closure state per exposure.
This is the principle the Picus Platform is built on, with Autonomous Penetration Testing, Exposure Validation and Breach and Attack Simulation all working off shared evidence.
Against a real clock: a critical CVE drops.
- Hour one, it’s enriched with threat intelligence.
- Hour two, affected assets, criticality, and surrounding controls are mapped.
- Hour four, exploitability is assessed on every affected asset, attack paths are confirmed where live testing is safe, and controls have been tested against the relevant techniques.
- Hour six, low-risk mitigations are deployed, the rest have been routed to owners, and every fix has been revalidated before closure.
Every change is validated. Every exposure is proven. Within hours. Request your demo and see one CVE run through all three methods on one shared findings model.
See it run, live
At The Validation Summit 26, hosted by Hacker Valley’s Ron Eddings, October 14 at 1 PM ET and October 15 at 11 AM BST, we’ll run this model live in the product. Our CTO Volkan Erturk will take an emerging CVE from trigger to validated exposure to a confirmed fix in hours, with the right method selected as conditions change. Additionally, Mikko Hyppönen will open with what changed after Mythos. Security leaders from Chanel, Atlassian, and Kraft Heinz will discuss how they’re preparing.
Two hours. Registration is free. See the workflow run live for yourself.
Found this article interesting? This article is a contributed piece from one of our valued partners. Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.

