Architect & SRE Master Curriculum
stanley-n.com

AWS Solutions Architect - Professional & SRE Master Curriculum

Comprehensive, first-principles curriculum designed for passing the AWS Certified Solutions Architect – Professional (SAP-C02) exam and building production-grade Site Reliability Engineering (SRE) / DevOps infrastructure on AWS.

1. 14-Week Architectural Roadmap

Exam Strategy Career Transformation

The Bridge from Developer to Solutions Architect / SRE

Associate certifications test individual service capabilities. The Solutions Architect – Professional tests multi-constraint optimization: balancing security guardrails, regional latency, blast-radius isolation, and RTO/RPO requirements while minimizing operational overhead.

Phase Duration Core Domain Focus SRE & DevOps Deliverable
Phase 1 Weeks 1–3 Multi-Account Governance, Control Tower, SCPs & Hybrid Networking Hub-and-Spoke Transit Gateway with inspection VPC & baseline SCPs
Phase 2 Weeks 4–6 Compute (ECS/EKS/Lambda), Storage Tiers & Global Databases Zero-downtime Blue/Green container deployments & Aurora Global DB
Phase 3 Weeks 7–9 SRE Focus: Observability, Resilience, Chaos & Disaster Recovery CloudWatch Synthetics SLOs, Chaos Game Days via AWS FIS, Route 53 ARC
Phase 4 Weeks 10–12 Enterprise Migration, Hybrid Cloud & Cost Optimization AWS Application Migration Service (MGN) & DMS CDC pipeline cutovers
Phase 5 Weeks 13–14 75-Question Timed Mock Exams, Distractor Analysis & Exam Review Pass 5 full-length Tutorials Dojo mock exams with >85% consistently

2. Day-by-Day Master Syllabus (70 Days)

Daily structured 1.5–2 hour study routine. Each day includes Core Architecture Reading (45m), Hands-on Practice (45m), and Scenario Question Drills (20m):

Day Topic Focus Week Phase
Day 1 Multi-Account Architecture & AWS Organizations Week 1 Multi-Account Governance & Networking
Day 2 Service Control Policy (SCP) Mechanics Week 1 Multi-Account Governance & Networking
Day 3 AWS Control Tower & Account Factory Week 1 Multi-Account Governance & Networking
Day 4 IAM Identity Center (AWS SSO) & Federation Week 1 Multi-Account Governance & Networking
Day 5 IAM Permission Boundaries & STS Week 1 Multi-Account Governance & Networking
Day 6 Resource Access Manager (RAM) & VPC Sharing Week 1 Multi-Account Governance & Networking
Day 7 Week 1 Synthesis & Multi-Account Drill Week 1 Multi-Account Governance & Networking
Day 8 Transit Gateway (TGW) Core Routing Week 2 Multi-Account Governance & Networking
Day 9 Transit Gateway Peering & Multi-Region Week 2 Multi-Account Governance & Networking
Day 10 AWS Network Firewall & Centralized Egress Week 2 Multi-Account Governance & Networking
Day 11 Direct Connect (DX) & DX Gateway Week 2 Multi-Account Governance & Networking
Day 12 AWS Route 53 Resolver (Hybrid DNS) Week 2 Multi-Account Governance & Networking
Day 13 AWS PrivateLink & VPC Peering Comparison Week 2 Multi-Account Governance & Networking
Day 14 Week 2 Synthesis & Networking Lab Exam Week 2 Multi-Account Governance & Networking
Day 15 Amazon ECS Fargate Architecture & Task Definitions Week 3 Compute, Containers & Global Databases
Day 16 Amazon EKS Architecture & Managed Node Groups vs Fargate Week 3 Compute, Containers & Global Databases
Day 17 Auto Scaling Groups — Lifecycle Hooks & Warm Pools Week 3 Compute, Containers & Global Databases
Day 18 Application Load Balancer — Advanced Routing & Target Groups Week 3 Compute, Containers & Global Databases
Day 19 AWS Lambda Advanced — Concurrency, VPC Networking & Event Source Mapping Week 3 Compute, Containers & Global Databases
Day 20 Serverless Architectures — Step Functions & API Gateway Week 3 Compute, Containers & Global Databases
Day 21 Week Synthesis — Compute Decision Matrix & Scenario Drills Week 3 Compute, Containers & Global Databases
Day 22 Amazon Aurora Architecture & Aurora Global Database Week 4 Compute, Containers & Global Databases
Day 23 Amazon RDS Multi-AZ vs Read Replicas & Blue/Green Deployments Week 4 Compute, Containers & Global Databases
Day 24 Amazon DynamoDB Deep Dive — Partitioning & Capacity Modes Week 4 Compute, Containers & Global Databases
Day 25 DynamoDB Global Tables & DAX Caching Week 4 Compute, Containers & Global Databases
Day 26 Amazon S3 Storage Classes, Lifecycle & Replication Week 4 Compute, Containers & Global Databases
Day 27 Amazon ElastiCache — Redis vs Memcached Architectures Week 4 Compute, Containers & Global Databases
Day 28 Week Synthesis — Database Selection Matrix & Scenario Drills Week 4 Compute, Containers & Global Databases
Day 29 CloudWatch Metrics, Alarms & Composite Alarms Week 5 SRE Observability, Resilience & DR
Day 30 CloudWatch Synthetics & Canaries for SLO Monitoring Week 5 SRE Observability, Resilience & DR
Day 31 AWS X-Ray Distributed Tracing Week 5 SRE Observability, Resilience & DR
Day 32 Amazon CloudWatch Logs Insights & Centralized Logging Week 5 SRE Observability, Resilience & DR
Day 33 AWS Fault Injection Service (FIS) — Chaos Engineering Fundamentals Week 5 SRE Observability, Resilience & DR
Day 34 Designing Game Days & Chaos Experiments Week 5 SRE Observability, Resilience & DR
Day 35 Multi-Region Active-Active Architectures Week 6 SRE Observability, Resilience & DR
Day 36 Multi-Region Active-Passive & Pilot Light Patterns Week 6 SRE Observability, Resilience & DR
Day 37 Route 53 Application Recovery Controller (ARC) Week 6 SRE Observability, Resilience & DR
Day 38 Route 53 Routing Policies & Health Checks Week 6 SRE Observability, Resilience & DR
Day 39 RTO/RPO Design Patterns & Backup Strategies (AWS Backup) Week 6 SRE Observability, Resilience & DR
Day 40 Well-Architected Reliability Pillar Deep Dive Week 6 SRE Observability, Resilience & DR
Day 41 Well-Architected Operational Excellence Pillar Week 6 SRE Observability, Resilience & DR
Day 42 Week Synthesis — SRE & DR Scenario Drills Week 6 SRE Observability, Resilience & DR
Day 43 Migration Strategy Framework — The 7 Rs & Migration Hub Week 7 Migration, Hybrid & Cost Optimization
Day 44 AWS Application Discovery Service — Assessment & Portfolio Analysis Week 7 Migration, Hybrid & Cost Optimization
Day 45 AWS Application Migration Service (MGN) — Lift-and-Shift Rehost Week 7 Migration, Hybrid & Cost Optimization
Day 46 AWS Schema Conversion Tool (SCT) — Heterogeneous Schema Conversion Week 7 Migration, Hybrid & Cost Optimization
Day 47 AWS DataSync — Hybrid File & Object Transfer Week 7 Migration, Hybrid & Cost Optimization
Day 48 AWS Storage Gateway — File, Volume & Tape Gateways Week 7 Migration, Hybrid & Cost Optimization
Day 49 AWS Snow Family — Snowball Edge, Snowcone & Snowmobile Week 7 Migration, Hybrid & Cost Optimization
Day 50 Hybrid Cloud Compute — Outposts, Local Zones & Wavelength Week 7 Migration, Hybrid & Cost Optimization
Day 51 VMware Cloud on AWS — Relocate Strategy Week 8 Migration, Hybrid & Cost Optimization
Day 52 Direct Connect Capacity Planning for Large-Scale Migration Week 8 Migration, Hybrid & Cost Optimization
Day 53 Cost Optimization — Savings Plans, RIs & Spot Fleet Strategy Week 8 Migration, Hybrid & Cost Optimization
Day 54 AWS Compute Optimizer, Cost Explorer & Trusted Advisor Week 8 Migration, Hybrid & Cost Optimization
Day 55 FinOps on AWS — Tagging, Budgets & Cost Allocation Reports Week 8 Migration, Hybrid & Cost Optimization
Day 56 AWS Database Migration Service (DMS) — Homogeneous & Heterogeneous Migration with CDC Week 8 Migration, Hybrid & Cost Optimization
Day 57 Full-Length Mock Exam 1 — 75 Questions Timed Week 8 Practice Exams & Exam Technique
Day 58 Mock Exam 1 Distractor Analysis & Weak Domain Review Week 9 Practice Exams & Exam Technique
Day 59 Deep Review — Domain 1 Weak Areas (Org/Governance/Networking) Week 9 Practice Exams & Exam Technique
Day 60 Deep Review — Domain 2 Weak Areas (Resilient Architectures) Week 9 Practice Exams & Exam Technique
Day 61 Full-Length Mock Exam 2 — 75 Questions Timed Week 9 Practice Exams & Exam Technique
Day 62 Mock Exam 2 Distractor Analysis & Weak Domain Review Week 9 Practice Exams & Exam Technique
Day 63 Deep Review — Domain 3 Weak Areas (Migration & Modernization) Week 9 Practice Exams & Exam Technique
Day 64 Deep Review — Domain 4 Weak Areas (Cost Control) Week 9 Practice Exams & Exam Technique
Day 65 Full-Length Mock Exam 3 — 75 Questions Timed Week 9 Practice Exams & Exam Technique
Day 66 Mock Exam 3 Distractor Analysis & Timing/Pacing Strategy Week 10 Practice Exams & Exam Technique
Day 67 Full-Length Mock Exam 4 — 75 Questions Timed Week 10 Practice Exams & Exam Technique
Day 68 Mock Exam 4 Distractor Analysis & Flag-and-Review Technique Week 10 Practice Exams & Exam Technique
Day 69 Full-Length Mock Exam 5 — 75 Questions Timed Week 10 Practice Exams & Exam Technique
Day 70 Final Review — Exam Day Logistics, Pearson VUE Checklist & Confidence Drills Week 10 Practice Exams & Exam Technique

3. Phase 1: Enterprise Multi-Account & Advanced Networking

Production Blueprint

Enterprise Service Control Policy (SCP) Guardrails

In enterprise architectures, developers must have full administrative permissions inside their local sandboxes or development accounts, while the platform team enforces non-bypassable organizational guardrails at the OU level.

Critical SCP Rules for SAP-C02:

  • SCPs only filter permissions—they never grant access. A user must still be explicitly granted permission via an IAM policy.
  • The AWS Organizations Management (root) account is immune to SCPs.
  • When writing region-restricting SCPs, always exempt global services (IAM, Route 53, CloudFront, WAF, AWS Support) to prevent breaking account operations.
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyUnapprovedRegionsExemptingGlobals",
      "Effect": "Deny",
      "NotAction": [
        "cloudfront:*",
        "iam:*",
        "route53:*",
        "support:*",
        "wafv2:*",
        "health:*",
        "organizations:*"
      ],
      "Resource": "*",
      "Condition": {
        "StringNotEquals": {
          "aws:RequestedRegion": [
            "us-east-1",
            "us-west-2"
          ]
        }
      }
    },
    {
      "Sid": "ProtectSecurityServices",
      "Effect": "Deny",
      "Action": [
        "cloudtrail:StopLogging",
        "cloudtrail:DeleteTrail",
        "config:DeleteConfigRule",
        "config:StopConfigurationRecorder",
        "guardduty:DeleteDetector",
        "guardduty:DisassociateFromMasterAccount"
      ],
      "Resource": "*"
    },
    {
      "Sid": "PreventLeavingOrganization",
      "Effect": "Deny",
      "Action": "organizations:LeaveOrganization",
      "Resource": "*"
    }
  ]
}
Networking Matrix

Hybrid Connectivity Decision Matrix

Connection Type Bandwidth Latency / SLA Setup Time Primary Architectural Use Case
Site-to-Site VPN Up to 1.25 Gbps per tunnel Variable (Public Internet) Minutes Backup for Direct Connect, quick POC, low-throughput workloads
Direct Connect (Dedicated) 1 Gbps, 10 Gbps, 100 Gbps Ultra-low, deterministic Weeks/Months Massive data migration, real-time trading, compliance requiring private physical wire
Direct Connect + Transit VIF Up to 100 Gbps Ultra-low Weeks/Months Direct Connect Gateway connected to Transit Gateways across thousands of VPCs
AWS PrivateLink Up to 100 Gbps per AZ Sub-millisecond Minutes Securely exposing services to customers or other VPCs with overlapping CIDR blocks

4. Phase 2: High-Availability Compute, Storage & Databases

Compute Architecture

ECS Fargate & EKS High-Availability Patterns

  • Fargate Spot Strategy: Run baseline production traffic on Fargate On-Demand (e.g. 30%), and burst peak traffic onto Fargate Spot (70%). Handle the 2-minute SIGTERM warning via EventBridge to drain connections cleanly.
  • Auto Scaling Lifecycle Hooks: When an EC2 instance in an ASG receives a termination signal, a lifecycle hook moves it to Terminating:Wait. A Lambda function triggers container connection draining and flushes local logs to S3 before signaling CONTINUE.
Global Data Stores

Aurora Global Database vs. DynamoDB Global Tables

Feature Amazon Aurora Global Database Amazon DynamoDB Global Tables
Replication Model Storage-level physical replication (<1s latency) Two-way multi-active logical replication (Streams)
Write Topology Single-region primary writer; cross-region read-only replicas Multi-region Active-Active (write to any region)
Disaster Recovery RTO < 1 minute (managed failover / promotion) Near-zero (seamless failover to other active regions)
Conflict Resolution N/A (all writes route to primary cluster) Last-writer-wins based on timestamp

5. Phase 3: SRE Focus — Reliability, Observability & Disaster Recovery

SRE Core

Production Observability & Incident Automation Stack

True SRE on AWS goes beyond passive dashboards. It couples active monitoring with automated remediation:

  • SLI / SLO Engineering: Use CloudWatch Metric Math to divide successful canary requests by total requests over a rolling 30-day window to track your Error Budget.
  • CloudWatch Synthetics: Deploy Node.js/Python Puppeteer canaries that simulate user checkout flows every 1 minute. Alarms fire on multi-step synthetic failure before real customers report outages.
  • AWS Fault Injection Simulator (FIS): Conduct scheduled Chaos Engineering Game Days in staging:
    • Inject 80% packet loss on cross-AZ network traffic.
    • Force failover of the primary Aurora database cluster.
    • Throttle API Gateway downstream endpoints to verify circuit breakers (via Amazon AppConfig & AWS Lambda).
Disaster Recovery

Disaster Recovery (DR) Architectural Decision Matrix

Strategy RTO (Recovery Time) RPO (Recovery Point) Relative Cost Key Implementation Services
Backup & Restore Hours to Days Hours to 24h $ AWS Backup, S3 Cross-Region Replication (CRR), EBS snapshot copies
Pilot Light 10 to 30 Minutes Seconds to Minutes $$ Aurora cross-region read replica, pre-baked AMIs, CloudFormation/Terraform IaC
Warm Standby Minutes Sub-second to Seconds $$$ Scaled-down EC2/ECS cluster running in DR region, Route 53 health check failover
Active-Active Multi-Region Near-Zero Near-Zero $$$$$ DynamoDB Global Tables, Route 53 ARC, AWS Global Accelerator, CloudFront

6. Phase 5: SAP-C02 Scenario Drill Bank

Scenario Drill #1
Scenario: A global media platform streams video to millions of users. The platform runs on an Amazon EKS cluster in us-east-1. Due to regulatory requirements, user analytics data must be replicated to eu-west-1 within 10 seconds of creation. However, the analytics ingest workload must never experience backpressure or HTTP 500 errors during sudden 10x traffic spikes. What architecture meets these requirements with the LEAST operational overhead?
A) Write analytics events directly to an Amazon S3 bucket in us-east-1 and configure S3 Cross-Region Replication (CRR) with Replication Time Control (RTC).
B) Ingest events into an Amazon Kinesis Data Stream in us-east-1. Use AWS Lambda to process events and push to an Amazon DynamoDB Global Table replicated to eu-west-1.
C) Publish events to an Amazon SNS topic in us-east-1 with cross-region subscriptions to an SQS queue in eu-west-1, processed by worker nodes on EKS.
D) Write events directly to an Amazon Aurora MySQL database in us-east-1 with AWS Database Migration Service (DMS) replicating to eu-west-1.
Correct Answer: B.
Rationale: Kinesis Data Streams reliably buffers massive ingestion spikes without dropping events or generating HTTP 500 backpressure. DynamoDB Global Tables provides fully managed, multi-region replication with sub-second replication latency (well within the 10-second SLA). Option A (S3 CRR with RTC) guarantees replication within 15 minutes, not 10 seconds. Option D creates database connection bottlenecks under 10x spikes.
Scenario Drill #2
Scenario: An enterprise has 200 AWS accounts organized in AWS Organizations. Security auditors discover that developers have created public Amazon S3 buckets containing proprietary scripts. The CISO mandates that no user in any account (except the Security-Audit role in the Security account) can ever create a public S3 bucket or alter S3 Block Public Access settings. What solution enforces this with the LEAST administrative effort?
A) Deploy an AWS Config rule s3-bucket-public-read-prohibited in every account with automatic remediation via AWS Systems Manager Automation.
B) Attach a Service Control Policy (SCP) at the Organizations Root with an explicit Deny on s3:PutBucketPublicAccessBlock and s3:PutAccountPublicAccessBlock unless the caller identity is the Security-Audit role. Enable S3 Block Public Access at the Organization level.
C) Write an Amazon EventBridge rule in every account that triggers an AWS Lambda function to delete any newly created public S3 bucket.
D) Apply an IAM Permissions Boundary to all users across all accounts denying S3 public access actions.
Correct Answer: B.
Rationale: An SCP applied at the Organization Root (or target OUs) provides a centralized, preventative guardrail that cannot be bypassed by individual account administrators. Options A and C are detective and reactive (the public bucket is briefly exposed before remediation). Option D requires managing and maintaining boundaries on every IAM entity in 200 accounts.
Scenario Drill #3
Scenario: A retail enterprise operates an e-commerce application on Amazon EC2 across multiple Availability Zones. During flash sales, the database becomes saturated with read requests, causing checkout latency. 85% of these queries fetch product catalog details that change only once every 24 hours. The solutions architect must reduce database load and maintain sub-millisecond response times with the least cost and operational complexity.
A) Upgrade the primary database instance to the largest memory-optimized Amazon RDS instance class and enable multi-AZ.
B) Deploy an Amazon ElastiCache for Redis cluster in front of the database using a lazy-loading (cache-aside) pattern with a 24-hour TTL.
C) Add 5 Amazon RDS read replicas and distribute queries using an Application Load Balancer.
D) Migrate the relational database to Amazon DynamoDB and enable DynamoDB Accelerator (DAX).
Correct Answer: B.
Rationale: ElastiCache (Redis) offloads repetitive static/slowly-changing queries directly from memory, delivering sub-millisecond response times and drastically reducing RDS CPU utilization. Lazy loading ensures only requested catalog items populate the cache. Option A is significantly more expensive and does not scale horizontally. Option C still queries disk storage and adds replica lag. Option D requires a complete data model rewrite.