AWS Penetration Testing for Enterprises: What to Scope and How to Budget
AWS penetration testing is scoped by accounts, identities, and exposed workloads, not IP ranges. AWS lets customers test a defined list of services without prior approval, but it prohibits flooding, DNS attacks, and takeovers, and it requires approval for command-and-control and simulated events. Budget follows two drivers: account and IAM complexity for the cloud control plane, and the number of internet-facing applications for workload testing.
Key Takeaways
- AWS penetration testing is scoped by accounts, IAM principals, services, and internet-facing workloads, not IP ranges.
- AWS lets customers test a defined list of services without prior approval. Flooding, DNS attacks, and takeovers are prohibited, and command-and-control and simulated events need approval first.
- The most important scoping decision is where the tester starts: outside with no credentials, inside with a low-privilege identity, or with a read-only audit role.
- Multi-account AWS Organizations force a choice between testing a sample of accounts and testing the whole estate, and the cross-account trust between them is often the highest-risk finding.
- The breaches that matter most in AWS, including Capital One in 2019, chained an ordinary flaw into an over-permissioned identity.
- Cost has two drivers: account and IAM complexity for the cloud control plane, and the number of exposed applications and APIs for workload testing.
- A CSPM scan lists misconfigurations. A penetration test shows which of them chain into a real path to your data.
What AWS Penetration Testing Covers
AWS penetration testing evaluates the customer’s side of the shared responsibility model: IAM configuration, resource permissions, application workloads, data stores, and network controls inside the customer’s own VPCs. AWS calls this “security in the cloud.” AWS owns “security of the cloud,” the infrastructure underneath, and testing never extends into it.
The customer’s share also changes by service. On EC2, the customer owns the guest operating system, installed software, and security group rules, so a test can go deep into the host. On managed services like S3 or DynamoDB, AWS runs the operating system and platform, and the customer’s attack surface shrinks to data handling, encryption settings, and IAM permissions. A good test plan follows that line service by service instead of treating every resource the same way.
Some things are always out of bounds, however a test is scoped: AWS’s own infrastructure, the hypervisor layer, and anything belonging to another AWS customer. A well-scoped AWS engagement stays inside the customer’s own accounts and identities. That is a different boundary from a traditional network penetration test, which is drawn around IP ranges.
The AWS Penetration Testing Policy in Plain Terms
AWS lets customers test a defined list of services without asking first. Before early 2019, testers had to fill out a request form and often wait up to a week for approval. Today the AWS penetration testing policy sorts activity into five buckets.
| Category | What it covers | What to do |
| Permitted without approval | EC2 (including WAF, NAT gateways, and load balancers), RDS, Aurora, CloudFront, API Gateway, AppSync, Lambda and Lambda@Edge, Lightsail, Elastic Beanstalk, ECS, Fargate, OpenSearch, FSx, Transit Gateway, Global Accelerator, Bedrock AgentCore | Test within your own accounts |
| Prohibited | DNS zone walking, hijacking, or pharming through Route 53; DoS, DDoS, and simulated DoS; port, protocol, and request flooding; S3 bucket takeover; subdomain takeover | Leave out of scope |
| Needs prior approval | Any testing that includes command-and-control (C2) | Request approval from AWS first |
| Needs a Simulated Events form | Red, blue, or purple team exercises; simulated phishing; malware testing; volumetric testing | Submit the form for AWS review |
| Separate DDoS policy | DDoS simulation | Use a pre-approved AWS partner against Shield Advanced resources |
Two details trip up enterprise buyers. First, S3 and EKS are not on the permitted list, even though both show up in most AWS estates, and S3 bucket takeover is prohibited outright. Testers usually assess them through the customer’s own authorized access, by reviewing bucket policies and what each identity can reach, rather than by attacking the service. Second, many enterprises want a full red team exercise, and that needs the Simulated Events form even when every target service is on the permitted list.
The list also changes as AWS adds services. Check the current policy before you finalize any engagement.
How to Scope an AWS Penetration Test
AWS scope cannot be written as an IP range, and buyers who try to scope it that way underestimate the effort. The first question is not what to test but where the tester starts, because that decides which findings are even reachable.
| Starting position | What the tester begins with | What it surfaces |
| External | No credentials, public view only | Exposed apps and APIs, leaked keys, public storage, guessable endpoints |
| Assumed breach | A low-privilege identity, as if an attacker already got a foothold | Privilege escalation, lateral movement, what one weak identity can reach |
| Configuration review | A read-only audit role across the accounts | Trust relationships, permission sprawl, and gaps the outside-in views miss |
Most mature programs combine the assumed-breach and configuration-review positions, because the highest-value AWS findings, such as privilege escalation, cannot be exercised without a starting identity. Testing purely from the outside tells you little about what happens once someone is in.
After the starting position, four inputs define the rest of the scope.
- Access. How the tester gets in: a dedicated IAM role with an external ID, or a set of test identities the customer provisions for the engagement.
- Rules of engagement. Production or non-production, allowed testing hours, and whether the security team and detective controls like GuardDuty should be tuned to catch the activity or left blind to test detection.
- Environment mapping. The scope units below, counted per account.
- External integrations. Third-party roles with cross-account trust, which may need authorization from that third party before testing.
| Scope unit | What to count | Why it drives effort |
| AWS accounts | In-scope account IDs, including shared services | Each account is a separate trust boundary |
| Organizational structure | OUs, SCPs, delegated admin accounts | Determines cross-account attack paths |
| Regions | Active regions per account | Resources hide in unused regions |
| IAM principals | Users, roles, policies, federation sources | The primary attack surface in the control plane |
| Internet-facing workloads | Public apps, APIs, and endpoints | The primary attack surface for the applications themselves |
| Service inventory | EC2, S3, Lambda, ECS, EKS, RDS, and others in use | Service diversity drives technique breadth |
| Workload types | Container, serverless, traditional compute | Each needs a different testing approach |
Two of these rows carry most of the weight, and they pull in different directions. For the control plane, IAM principals predict effort far better than raw resource count: an account with fifty roles and complex federation takes longer to test than one with five roles and more compute. For the applications, the number of internet-facing apps and APIs is what matters, regardless of how simple the IAM setup is. A good scope sizes both, rather than reducing the whole engagement to one number.
Scoping Across AWS Organizations and Multi-Account Estates
Enterprise AWS estates are rarely one account. A landing-zone pattern can run dozens or hundreds of accounts under one Organization, tied together by cross-account role chains, shared-services accounts, and delegated administration. Testing across that structure raises a decision a single-account test never faces: test a representative sample of accounts, or the whole estate.
Testing every account in a two-hundred-account Organization is rarely practical. Testing only a handful risks missing the cross-account trust relationships that pose the greatest risk in a mature environment. The workable middle ground is to test a deliberate sample, chosen to include the accounts with elevated trust or sensitive workloads, plus a dedicated focus on the role chains that connect them.
This is also where a scan and a test divide the work cleanly. Automated configuration checks can cover every account cheaply, which keeps the inventory honest. Human testers then go deep on the sample and on the trust between accounts, which is where chained findings live. Continuous attack surface management between engagements helps too, since a landing zone that looked complete six months ago has usually grown by the next test.
The AWS Attack Paths a Good Test Looks For
A strong AWS test focuses on the paths by which AWS is actually compromised, not generic network techniques aimed at AWS-hosted infrastructure. The public breach record shows the same four categories again and again. Each one below is paired with a real incident and with what a test should verify.
IAM Privilege Escalation
Misconfigured policies and permission chains that let a low-privilege identity reach admin access are the most common and most important AWS finding. Rhino Security Labs catalogued 21 distinct IAM privilege-escalation methods back in 2018, and the underlying misconfigurations still appear in most engagements. A test should confirm whether any identity a tester can reach can grant itself more permissions, and remediation is about tightening specific policies rather than adding tools.
Cross-Account Trust Abuse
Over-broad trust policies and weak or missing external IDs turn a foothold in one account into access across the Organization. AWS documents this as the confused deputy problem and recommends sts:ExternalId, aws:SourceArn, and aws:SourceAccount as the fixes. That guidance is widely ignored: Datadog’s 2025 State of Cloud Security still found third-party integration roles that don’t enforce an external ID. In 2023, Datadog researchers also found over 500 IAM roles across more than 275 accounts whose GitHub Actions trust policies were missing subject validation, including one at the UK Government Digital Service. This is exactly the finding a single-account test, scoped without regard to the wider Organization, never surfaces.
Data Exposure
Storage and snapshot settings, plus secrets left in code or environment variables, remain among the most common findings, and one misconfigured permission can expose a lot of data. Uber’s 2016 breach began with an AWS access key in a private GitHub repository that opened an S3 datastore holding data on roughly 57 million riders and drivers, including 600,000 driver’s license numbers. The Codecov Bash Uploader compromise in 2021 harvested CI credentials and environment variables from customers’ pipelines for two months before it was caught. And Bishop Fox’s DEF CON 27 research on public EBS snapshots found passwords, SSH keys, TLS certificates, and API keys in volumes their owners never meant to expose. A test should check what is actually reachable, not just what a policy claims.
Workload Compromise Paths
Instance metadata exposure, container escape, and over-permissive task or node roles round out the list. A compromised workload with a broad role can often reach far more than it was ever meant to. Tesla’s 2018 cloud incident is the textbook case: RedLock found an unprotected Kubernetes console that exposed AWS credentials, which attackers used for cryptomining inside Tesla’s environment. AWS’s own response to this class of risk was IMDSv2, launched in November 2019 to add defense in depth against metadata theft. Adoption is still partial: Datadog’s 2025 report found IMDSv2 enforced on 49% of EC2 instances. A test should confirm whether workloads enforce IMDSv2 and whether their attached roles follow least privilege.
The pattern behind all four
The most-cited AWS breach ties these together. In 2019, Capital One suffered a breach affecting about 100 million people in the US and 6 million in Canada. The chain was ordinary: a misconfigured web application firewall with too many permissions let an attacker retrieve credentials and read data from S3. No single flaw was exotic. An over-permissioned identity turned one web vulnerability into a nine-figure data loss, and the OCC later fined Capital One $80 million for failing to assess cloud risk before migrating. That chaining, an ordinary flaw plus an over-broad identity, is what a good AWS test is built to find before someone else does.
Is a CSPM Scan the Same as an AWS Penetration Test?
No, and the difference is the whole reason to buy a test. A Cloud Security Posture Management (CSPM) scan checks your configuration against a baseline and produces a list of misconfigurations: this bucket allows public access, that role is over-permissive, this instance still allows IMDSv1. That list is useful, and a good program runs one continuously.
What a scan cannot tell you is which of those findings actually chain into a path to your data. A penetration test does. In the Capital One breach, each ingredient, a web-app flaw and an over-permissioned role, might have appeared on a scan as a moderate item. The severity only became clear when they combined into a route to 100 million records. A scanner sees the pieces; a tester proves whether the pieces connect.
The two are complementary, not competing. Use scanning for breadth, so every account is checked cheaply and often. Use human testing for depth, to validate whether the flagged issues are exploitable and to find the chains no baseline check is looking for. When a report says a finding is real, it should mean a tester reached it, not that a rule fired.
What Drives AWS Penetration Testing Cost
Cost moves with a specific set of variables. Knowing them lets a buyer sanity-check a quote against their own environment instead of reacting to a single number. The variables fall into the same two groups the scope did.
Control-plane effort scales with:
- Account count, since each account is a separate trust boundary to test.
- IAM principal count, since authorization testing scales with users, roles, and federation sources.
- Organization complexity, since cross-account role chains and a landing-zone review add work a single-account test never does.
Workload effort scales with:
- Internet-facing apps and APIs, the count of exposed surfaces that need active testing.
- Service diversity, since a broader mix of services requires broader technique coverage.
- Workload types, since container, serverless, and traditional compute each need a different approach.
Engagement structure then adjusts both:
- Whether a configuration baseline review runs alongside active testing.
- The retest policy, since confirming that IAM and configuration fixes actually closed the exposure adds real time.
- Whether a landing-zone review across the wider Organization is included.
Published market ranges vary widely because providers scope these differently from one proposal to the next. Treat any figure as a starting point and confirm it against your own account count, IAM footprint, and exposed workloads before you commit.
How to Budget an Enterprise AWS Test
Start by sizing the two drivers, not by asking for a single price. A useful first pass is to count in-scope accounts, estimate IAM principals across them, and list the internet-facing apps and APIs. Those three numbers tell a provider more than any asset inventory.
Consider a mid-size estate: one Organization, roughly 40 accounts with about 8 of them holding sensitive workloads or elevated trust, a few hundred IAM roles, and a dozen internet-facing applications. A sensible engagement there would not test all 40 accounts. It would test the 8 high-value accounts and the role chains connecting them, run a configuration review across the rest, and test the dozen exposed apps. Sizing it this way keeps effort tied to risk instead of to raw account count, and it gives the provider a clear basis for a quote.
Three strategies keep the budget matched to risk:
- Phase it. Test the highest-value accounts and exposed apps first, then widen coverage in later engagements as the program matures.
- Sample, then expand. Start with a representative set of account types and grow the sample as the estate grows, rather than paying to test everything at once.
- Split one-time from continuous. A point-in-time test proves a moment; continuous testing tracks an estate that changes weekly. Many enterprises run periodic deep tests on crown-jewel accounts and continuous coverage across the rest.
Budget for the retest, too. A test that finds an over-permissioned role is only half the value; confirming the fix actually closed the path is the other half, and it should be a line item, not an afterthought.
Finally, weigh the cost against the downside. IBM put the 2025 global average cost of a data breach at $4.44 million, and Capital One’s cloud breach drove an $80 million regulatory fine on top of remediation and litigation. Against numbers like those, the test is the cheap part.
What the AWS Penetration Testing Report Should Contain
An AWS report has to be specific enough for a cloud team to act on without a second meeting. Generic findings force that meeting; precise ones don’t. Look for:
- Findings mapped to exact accounts and resources, by account ID and ARN, not “an S3 bucket” or “a role.”
- Escalation paths shown as chains, from initial access to elevated privilege, so the reader sees how the pieces connect rather than a scatter of isolated issues.
- Specific remediation, the IAM policy or configuration change to make, not “apply least privilege.”
- A distinction between fix types. Some findings an engineer can close directly; others need a cloud architect to redesign part of the account structure. A strong report says which is which instead of treating every finding as equally actionable.
- Compliance mapping where it applies, tying findings to PCI DSS requirement 11.4, SOC 2, or the FedRAMP attack vectors such as tenant-to-tenant and tenant-to-management-plane, so the test does double duty for the audit.
- Retest evidence, confirming that the fixes actually closed each path.
- Two audiences. An auditor-ready summary alongside the technical detail an engineering team needs.
The test is only as good as what the customer can do with it afterward, and that is decided by the report.
Frequently Asked Questions
Not for the permitted service list, which covers EC2, RDS, Lambda, API Gateway, and others. AWS dropped the pre-approval requirement in 2019. Some activities still need approval first: any command-and-control testing, and red team exercises, simulated phishing, or malware testing, which require an AWS Simulated Events form.
AWS’s own infrastructure, the hypervisor layer, and any other customer’s resources are always off-limits, however the test is scoped. AWS also prohibits denial-of-service and simulated DoS, port and request flooding, Route 53 DNS attacks, and S3 or subdomain takeover. A well-scoped test stays inside your own accounts and identities.
By accounts, IAM principals, and the cross-account trust between them, not by IP range. In a large Organization, testers usually assess a deliberate sample of high-value accounts plus the role chains that connect them, rather than every account. A configuration review can cover the remaining accounts for breadth.
Cost tracks two drivers: account and IAM complexity for the cloud control plane, and the number of internet-facing apps and APIs for the workloads. Resource count matters far less. Budget by counting in-scope accounts, estimating IAM principals, and listing exposed applications, then confirm any published range against those numbers.
No. A CSPM scan checks configuration against a baseline and lists misconfigurations. A penetration test proves which of those actually chain into a real path to your data. Scanning gives breadth across every account; human testing gives depth and finds the chains no baseline rule is looking for. The two are complementary.


