Palo Alto Networks has turned techniques traditionally performed in one-off security tests into a continuous service. Launched on September 22, the Unit 42 Continuous Frontier AI Defense uses Anthropic, OpenAI and open-weight models to look for vulnerabilities, try to combine them into real attack paths and recommend fixes as applications and infrastructure change. The proposal puts pressure on an established corporate security practice: hiring pentests at defined intervals and accepting that, between one assessment and the next, the environment will keep changing.
The question is not whether traditional pentest disappears. The service's own characteristics indicate the opposite. What is starting to change is which part of the work needs to wait for the next human test.
A year of testing compressed into three weeks
Unit 42 says it developed the approach over six months, with US$ 17 million in research, and validated the method in more than 100 customer engagements. According to the company, its internal deployment found in three weeks an amount of exposures equivalent to that identified by approximately a year of traditional testing. The system also reportedly located 3.2 times more high- or critical-severity vulnerabilities per product and reduced by 51% the average time required for remediation.
In the customer environments evaluated, Palo Alto says it found exposures in 100% of cases, with 37% classified as high or critical. In third-party applications, two-thirds of validated exposures were not associated with a known CVE.
That last data point helps explain why the comparison with vulnerability scanners is limited. A CVE describes a known vulnerability. An attack path can arise from the combination of minor flaws, inadequate permissions, application logic and configurations that individually do not appear as a cataloged vulnerability.
The service tries precisely to bridge that gap. After an initial scan of the environment, agents continue running tests as applications, APIs, identities, code repositories, cloud and other assets change. Human experts at Unit 42 validate the results and attack chains before remediation.
The problem with periodic pentest is the interval between two tests
The classic model did not arise from a lack of interest in continuous testing. There is an economic and operational limitation.
NIST's security testing guide notes that penetration tests can be expensive and pose a risk to production systems. For this reason, the agency considered that, depending on the organization, an annual test could be sufficient, complemented by less intensive activities and scanners run more frequently.
The recommendation is from 2008. Since then, corporate infrastructure has become much more dynamic, with cloud, APIs, continuously delivered software and external dependencies capable of modifying the attack surface without waiting for the next audit cycle.
NIST itself already treats continuous monitoring as a separate component of risk management, intended to maintain permanent visibility over assets, vulnerabilities and the effectiveness of controls.
Offensive agents add a new layer to this concept: not only continuously observing whether a vulnerability exists, but trying to find out whether it can actually be exploited and combined with other flaws.
If this approach works at scale, the most relevant consequence will be reducing the period during which a new exposure remains unknown simply because the next pentest has not yet started.
The attacker's speed makes that interval more costly
The incentive to reduce that window is increasing.
Mandiant data show that exploits continued to be the most frequent initial vector of intrusions investigated in 2025, accounting for 32% of cases. The report also recorded a drop in the median time between first access and the transfer of that access to a second criminal group, from more than eight hours in 2022 to just 22 seconds in 2025.
AI adds automation capability to this scenario. In 2026, the Google Threat Intelligence Group documented the first publicly confirmed case in which a criminal used AI to assist in the discovery and transformation of a zero-day vulnerability into an exploit intended for a mass exploitation campaign.
This does not mean that autonomous AI attacks are already responsible for the majority of intrusions. Unit 42 itself states that it has not yet observed a fundamental change in the techniques employed by attackers: so far, AI appears mainly as an efficiency multiplier for already known methods. Mandiant reached a similar conclusion when analyzing 2025 incidents, stressing that traditional human and systemic failures remain behind most successful intrusions.
The change, therefore, lies first in the speed and cost of searching for opportunities, not necessarily in the creation of a completely new class of attacks.
Pentest does not disappear, but it may change its role
Automating discovery and exploitation also does not eliminate the need for specialists.
The very architecture presented by Unit 42 keeps humans in the validation process. There is another revealing limitation: according to information obtained by Axios, in the company's tests no single model managed to locate more than 40% of vulnerabilities in complex environments. The solution uses multiple models precisely because they have different capabilities and blind spots.
This points to a division of labor that is more likely than the simple replacement of pentest.
Agents can repeatedly perform tasks that are scalable by software: exploit applications, test configurations, search for combinations of flaws and redo attacks after changes in the environment.
Human pentesters remain relevant in situations that depend on context, adversarial creativity, business logic, social engineering, judgment about impact and tests that cannot be permanently executed against critical systems.
The result may be a reversal of the current logic. Instead of periodic pentest being the main moment when a company tries to discover how it would be attacked, technical discovery starts to occur continuously and deeper human assessments become a complementary layer of validation and exploration of specific scenarios.
It is this shift that Palo Alto's launch makes more important to follow. The decisive indicator will not be how many vulnerabilities an agent can list, but how many valid attack paths it can find without generating excessive noise and how much time companies can actually save between discovery and correction.
If metrics like those reported by Unit 42 are reproduced in larger environments and by other vendors, the discussion about pentest tends to stop being just “how many times a year should we test?” and start being “which parts of our infrastructure can still go without being continuously tested?”



