Home Blog Penetration Testing Frequency

Penetration Testing Frequency: Why Annual Tests Fail in Modern Environments

Penetration Testing Frequency: Why Annual Tests Fail in Modern Environments

Consider a familiar scenario. Your annual penetration test is complete, and the report has finally arrived. There are dozens of findings in one PDF, ranked by severity. The engineering team starts going through them. Some are easy to reproduce and fix. Others need investigation: where does this code live, is the endpoint still in use, has the issue already been fixed somewhere else? A few critical findings get attention immediately. The rest are added to the backlog.

By the time the backlog is cleared, the application has changed again.

The test measured one version of the system. The system ships a new one every week.

Engineering shortened its feedback loop. Security didn’t.

For a decade, engineering has worked to close the gap between making a change and finding out whether it was any good. Nightly builds became CI. Release branches gave way to trunk-based development. Big-bang releases became feature flags and canaries. Manual QA gates became test suites that run on every commit. Same idea each time: shrink the batch, catch the defect near the point where it was introduced.

Security testing kept the old cadence. The 2025 DORA data puts about 24% of organizations below monthly deployment — meaning roughly three-quarters ship at least monthly, and 16% ship on demand. A team on a daily cadence puts around 250 releases into production between annual assessments. One of those 250 was tested.

There’s a number worth tracking that almost nobody tracks: how many production releases happen between scope freeze and the next test? That is the size of the untested gap.

Nobody would accept this from a functional test suite. Security gets a pass because penetration testing is usually managed as a compliance line item rather than as part of the engineering feedback loop.

Scope Freezes Before the Risky Code Exists

Scope is negotiated weeks ahead of the engagement, then locked. Code that ships after the freeze isn’t tested. Systems changed or retired before the engagement may still be in the report.

The untested part isn’t a random sample. It skews toward the newest code — the feature that shipped last week, the integration added under deadline pressure, the service with no operational history behind it. Those are the parts most likely to be wrong.

The Attack Surface Moved Past the Application

A modern application depends on far more than its login page and public API.

CI/CD pipelines hold deployment credentials, repository access, package-publishing rights and signing keys, and they execute third-party code on every run. Cloud identities carry permissions granted years ago by someone who has since left. In most environments the delivery pipeline is one of the most privileged systems in the organization, and it is routinely out of scope for the application penetration test.

Attackers have noticed. In September 2025, CISA issued an alert on Shai-Hulud, a self-propagating compromise of hundreds of npm packages that harvested npm tokens, GitHub PATs and cloud keys, then republished itself through the maintainers’ own accounts. The November wave was substantially larger — 25,000+ malicious repositories — and moved its payload to preinstall, so no human interaction was required.

A month before the first wave, a separate group (tracked as UNC6395) used stolen OAuth tokens from the Salesloft Drift integration to reach the Salesforce tenants of hundreds of companies. What they pulled out was support tickets containing plaintext AWS keys, VPN credentials and Snowflake tokens.

Neither attack touched a login form. Both ran through the supply chain and through identities.

This doesn’t mean every penetration test should cover the whole development environment. It means somebody has to decide, explicitly, which parts of it are tested and which are not. An application test is not a test of the systems that build, deploy and connect that application.

AI-Assisted Development Raises the Volume

The problem predates coding agents. What changed is throughput: more code is produced and modified, without a matching increase in review capacity.

Two effects matter. A flawed pattern now propagates faster — an authorization mistake that used to sit in one endpoint becomes the house style across a service. And coding agents are becoming identities in their own right: repository access, CI permissions, tokens, access to external tools. Their blast radius is a function of how those permissions were configured, which makes agent privilege management part of the attack surface rather than an AI tooling question.

What Penetration Testing Is Actually Good At

None of this is an argument for less penetration testing. It’s an argument against treating it as a calendar event whose output is a PDF.

Scanners don’t chain three low-severity findings into account takeover. They don’t find business-logic flaws where every individual request is valid and only the sequence is wrong. They don’t challenge an assumption in a threat model that looked reasonable to the people who wrote it. And they don’t provide the independent outside assessment that customers, partners and regulators ask for.

The craft holds up. The scheduling doesn’t.

How Often Should You Actually Test? Three Cadences, Not One

Continuous, at the base. Dependency and supply-chain scanning, secret detection, static analysis, cloud configuration and permission drift, and an accurate inventory of the external attack surface. You can’t test what you don’t know exists.

Triggered by change, in the middle. A deep review should follow a change in the risk profile, not the passage of twelve months. Triggers worth writing down:

  • a new or reworked authorization model
  • an internal API becoming external
  • changes to tenant-isolation logic
  • a new integration with broad permissions
  • a new authentication mechanism
  • a major migration
  • an acquisition, or integrating acquired infrastructure
  • material change to CI/CD or identity architecture

Adversarial, at the top. Annual in some organizations, tied to architectural or business events in others. What matters is the question it answers. Most engagements ask what an attacker could exploit. Add the second one: if someone actually tried, how long before we noticed? A patched vulnerability closes one known path; detection covers the paths nobody wrote a test for.

Metrics Worth Putting Next to Your DORA Numbers

Counting findings tells you about the last engagement, not about whether the program is improving. Four that do:

Recurrence rate by defect class. The same class of bug coming back means findings are being closed as tickets instead of fed back into engineering controls.

Share of findings caught before production. Rising means the controls are moving closer to where code is written.

Time to remediate by severity. Not how many criticals existed, but how long they stayed exploitable.

Time to detect during adversarial testing. How long testers operate before defenders notice. This one is uncomfortable by design, which is most of its value.

The Compliance Objection Is Weaker Than It Looks

The standard defense of the annual cycle is that an auditor requires it. The requirement is usually broader than that. PCI DSS 4.0 requires internal and external penetration testing at least every 12 months and after any significant infrastructure or application change (11.4.2, 11.4.3). The change trigger is already in the standard. Most programs just never operationalized it.

PCI DSS is not the only framework where this applies, but it is the one that says so explicitly.

Change-triggered testing doesn’t mean dropping the annual assessment. In practice it means keeping it and adding focused testing when the risk profile moves.

The Bottom Line

The annual penetration test isn’t broken. It’s being asked to stand in for security testing of a system that ships every week. Its strengths are depth, independence and human judgment. Its weakness is timing.

The fix isn’t more tests or fewer. It’s a different trigger: machines test continuously, people test when the risk profile moves, and adversarial exercises measure how long it takes you to notice.

Four steps to start, no new budget required:

  1. Count the gap. Count the production releases between the last scope freeze and the next test.
  2. Write the triggers down. Attach them to a process that already runs, such as architecture review, change approval or the RFC template.
  3. Make scope explicit. List what the next engagement won’t cover, such as pipelines, identities and agent permissions, and who accepted that risk.
  4. Add time-to-detect to the next adversarial test. The first answer will be uncomfortable. That’s your baseline.

Engineering shortened its feedback loop. Security testing has to do the same. Keep the annual test, but stop treating it as the program.

FAQ

As a floor, at least once every 12 months — but frequency should follow risk, not the calendar. Continuous automated scanning covers the base layer, testing triggered by significant changes (new auth model, new integration, CI/CD changes) covers the middle, and periodic independent adversarial testing covers the top. Annual alone is a minimum, not a program.

PCI DSS 4.0 requires internal and external penetration testing at least once every 12 months (Requirements 11.4.2 and 11.4.3), and again after any significant infrastructure or application change. Both conditions apply — the standard already includes a change trigger, even though most programs only track the annual one.

For a static environment, an annual test may be sufficient. For any organization deploying more than a handful of times a year, it isn’t: the gap between scope freeze and the next test is where new features, new integrations, and new attack surface accumulate untested. Annual testing should be the baseline layer of a program, not the whole program.

A new or reworked authorization model, an internal API becoming external, changes to tenant-isolation logic, a new integration with broad permissions, a new authentication mechanism, a major migration, an acquisition, or a material change to CI/CD or identity architecture.

Vulnerability scanning is inexpensive enough to run continuously and should. Penetration testing requires human judgment to chain findings, test business logic, and validate assumptions — so it is scheduled around triggers and independent assessment rather than run continuously in the same way.

Not Sure Where Your Testing Gap Actually Is?

The gap could be at either layer. Maybe there’s no continuous visibility between engagements — nobody would notice a new exposed service or a permission drift until next year’s report. Or maybe the last authorization change, new integration, or CI/CD overhaul shipped with no testing at all, because it didn’t line up with the calendar.

CodeIT covers both. Continuous Threat Exposure Management (CTEM) handles the base layer — ongoing assessment of your attack surface, so nothing sits unmonitored between annual tests. Penetration Testing engagements handle the middle layer — scoped to the change that actually happened, not to a date on the calendar.

About author
Photo of Mykhailo Momot
Full Stack Engineer
As a Full Stack Engineer at CodeIT, Mykhailo bridges robust backend engineering with modern frontend technologies. Bringing 5+ years of hands-on experience across Node.js, PHP, and Angular, he routinely drives complex third-party integrations and system refactoring. Mykhailo is committed to code quality, automated testing, and creating seamless digital experiences.

Business First
Code Next
Let’s talk

    By clicking the “Send” button I confirm, that I have read and agree to the Privacy Policy.