Skip to content

Aws Advanced

← Back to all decks

15 cards — 🟢 1 easy | 🟡 13 medium | 🔴 1 hard

🟢 Easy (1)

1. How do cost allocation tags differ from regular resource tags?

Show answer Regular tags: key-value metadata on AWS resources for organization, access control (IAM conditions), automation.

Cost allocation tags: regular tags activated in Billing Console for cost tracking.

Two types:
1. AWS-generated: aws:createdBy, aws:cloudformation:stack-name — activate in Billing
2. User-defined: any tag you create — must activate in Billing

After activation:
- Tags appear in Cost Explorer and Cost & Usage Report
- Can filter/group costs by tag
- Takes 24 hours to appear in billing data
- Not retroactive — only applies to future usage

Best practices:
- Mandate tags via SCP or AWS Config rules
- Standard tags: Environment, Team, Project, CostCenter
- Use Tag Editor for bulk tagging
- Enforce with aws:RequestTag conditions in IAM

Tags only affect billing visibility, not actual charges.

🟡 Medium (13)

1. How does RDS Multi-AZ failover work?

Show answer RDS Multi-AZ maintains a synchronous standby replica in a different AZ.

Failover triggers: AZ outage, instance failure, OS patching, manual reboot with failover.

Process:
1. DNS endpoint (CNAME) flips to standby (60-120 seconds)
2. Standby promoted to primary
3. Old primary becomes new standby when recovered

Connection handling:
- Existing connections are dropped — apps must reconnect
- Use DNS caching TTL of 5s or less
- Connection poolers should validate connections

Not for read scaling — standby is not readable (use Read Replicas for that). Aurora Multi-AZ is faster (typically <30s) because it uses shared storage. The endpoint doesn't change — applications just need retry logic.

2. Explain S3 lifecycle policies and transition timing

Show answer Lifecycle policies automate transitioning objects between storage classes or deleting them.

Transition rules (minimum days from creation):
S3 Standard -> S3 IA: 30 days minimum
S3 Standard -> S3 Glacier IR: 90 days minimum
S3 IA -> Glacier: 30 days after IA transition
Any -> Glacier Deep Archive: 90 days minimum

Key rules:
- Minimum object size: 128KB for IA transitions (smaller objects stay in Standard)
- Can't transition backwards (Glacier -> Standard requires restore)
- Versioning: can target current/noncurrent versions separately
- Expiration: delete objects after N days

Example policy: Transition to IA at 30 days, Glacier at 90 days, delete at 365 days. Cost savings of 40-80% for infrequently accessed data.

3. What are CloudFormation intrinsic functions?

Show answer Built-in functions for dynamic value resolution in templates:

!Ref LogicalName — returns resource ID or parameter value
!GetAtt Resource.Attribute — get resource attribute (e.g., ARN, DNS name)
!Sub 'Hello ${AWS::AccountId}' — string substitution
!Join ['-', [a, b, c]] — join with delimiter: 'a-b-c'
!Select [1, [a, b, c]] — pick index: 'b'
!Split [',', 'a,b,c'] — split string into list
!If [condition, true_val, false_val] — conditional
!ImportValue ExportName — cross-stack references

Common pattern:
BucketArn: !GetAtt MyBucket.Arn
BucketUrl: !Sub 'https://${MyBucket}.s3.amazonaws.com'

Short form (!) is YAML-specific. JSON uses { 'Fn::Ref': ... }. !Ref is the most commonly used function.

4. Explain Lambda concurrency: reserved, provisioned, and burst

Show answer Lambda concurrency = number of simultaneous function executions.

Types:
- Unreserved: shared pool (default 1000 per account per region)
- Reserved: guaranteed concurrency for a function (subtracted from account pool)
- Provisioned: pre-initialized instances (no cold starts, costs more)

Burst behavior:
- Initial burst: 500-3000 (region-dependent)
- After burst: +500 concurrent executions per minute

Throttling: returns 429 (sync) or retries (async, 2 attempts).

Cold start factors: runtime (Java worst, Python/Node best), VPC (adds ENI setup), package size. Provisioned concurrency eliminates cold starts but costs ~$0.015/GB-hour.

Best practices: right-size reserved concurrency, use provisioned for latency-sensitive paths, set account-level concurrency limits to prevent runaway costs.

5. ALB vs NLB: when to use which?

Show answer ALB (Application Load Balancer) — Layer 7 (HTTP/HTTPS):
- Path/host/header-based routing
- WebSocket support
- Lambda targets
- Authentication integration (OIDC, Cognito)
- WAF integration
- Slower (HTTP parsing overhead)
- ~$0.023/hour + LCU charges

NLB (Network Load Balancer) — Layer 4 (TCP/UDP/TLS):
- Ultra-low latency (~100us vs ~400us for ALB)
- Static IP / Elastic IP support
- Preserves source IP by default
- Millions of requests per second
- TCP passthrough (no HTTP parsing)
- PrivateLink support
- ~$0.023/hour + NLCU charges

Use ALB for: HTTP APIs, microservices, path-based routing, Lambda backends.
Use NLB for: extreme performance, TCP/UDP protocols, static IPs, gaming, IoT, PrivateLink endpoints.

6. How do AWS Organizations and Service Control Policies work?

Show answer Organizations manages multiple AWS accounts in a hierarchy of OUs (Organizational Units).

Service Control Policies (SCPs):
- Permission guardrails applied to OUs or accounts
- Restrict what services/actions are ALLOWED (deny list or allow list)
- Don't grant permissions — they filter existing IAM permissions
- Don't affect the management account

Example SCP (prevent leaving org):
{
\"Effect\": \"Deny\",
\"Action\": \"organizations:LeaveOrganization\",
\"Resource\": \"*\"
}

Key behaviors:
- SCPs are inherited down the OU tree
- Effective permissions = IAM policies INTERSECTED with SCP chain
- Full access SCP attached by default — removing it blocks everything
- Cannot restrict root user in management account

Pattern: deny-list approach (allow all, deny specific risky actions) is easier to manage than allow-list.

7. Explain VPC peering vs Transit Gateway

Show answer VPC Peering:
- 1:1 connection between two VPCs
- Non-transitive (A-B and B-C doesn't mean A-C)
- No bandwidth bottleneck (uses AWS backbone)
- No cost for data transfer in same AZ
- Cross-region and cross-account supported
- Limit: 125 peering connections per VPC
- No overlapping CIDRs allowed

Transit Gateway:
- Hub-and-spoke model for connecting multiple VPCs
- Transitive routing (A-B-C works through TGW)
- Supports VPN and Direct Connect attachments
- Route tables for network segmentation
- Bandwidth: up to 50 Gbps per attachment
- ~$0.05/hour + $0.02/GB data processing

Use peering for: few VPCs, simple connectivity, cost optimization.
Use TGW for: many VPCs (>10), centralized routing, hybrid connectivity, network segmentation at scale.

8. How does AWS Config differ from CloudTrail?

Show answer CloudTrail — WHO did WHAT and WHEN:
- Records API calls (management + data events)
- Shows who called which API, from where, when
- Used for security investigation and audit trails
- Delivers to S3, CloudWatch Logs, EventBridge

AWS Config — WHAT is the current/historical STATE:
- Tracks resource configurations over time
- Shows how a resource was configured at any point
- Evaluates compliance rules (Config Rules)
- Can auto-remediate with SSM Automation

Example: 'Who deleted the security group?' = CloudTrail
'What did the security group look like before deletion?' = Config
'Are all security groups compliant with our rules?' = Config Rules

Use together: CloudTrail for the action trail, Config for the configuration timeline and compliance posture.

9. What are S3 access points and when should you use them?

Show answer S3 Access Points are named network endpoints with distinct permissions policies, simplifying access management for shared buckets.

Without access points: one complex bucket policy managing all access patterns.
With access points: each team/application gets its own access point with its own policy.

# Create:
aws s3control create-access-point \\
--name analytics-team \\
--bucket shared-data \\
--vpc-configuration VpcId=vpc-123

Features:
- Each access point has its own DNS name
- Can restrict to specific VPC (no internet access)
- Own IAM policy (simpler than bucket policy)
- Block public access settings per access point
- S3 Object Lambda access points for data transformation

Use when: multiple teams share a bucket, complex access patterns, need VPC-scoped access, per-application permissions. Limit: 10,000 access points per account per region.

10. Explain ECS task definitions vs EKS pod specs

Show answer Both define container workloads but use different models:

ECS Task Definition:
- JSON document defining 1+ containers
- Specifies: image, CPU/memory, ports, env vars, logging
- Launch type: Fargate (serverless) or EC2
- Task role (IAM) for AWS API access
- Execution role for ECR pulls and logging
- Service = long-running tasks with desired count

EKS Pod Spec:
- YAML manifest with Kubernetes API objects
- Pod: 1+ containers sharing network/storage
- Deployment/StatefulSet for management
- Service accounts + IRSA for AWS API access
- More flexible: init containers, sidecars, volumes
- Richer networking: NetworkPolicy, Ingress

Choose ECS for: simpler workloads, AWS-native, less operational overhead.
Choose EKS for: Kubernetes expertise, multi-cloud portability, complex networking, existing K8s workloads.

11. What is Kinesis vs SQS vs SNS — when to use which?

Show answer SNS (Simple Notification Service) — pub/sub:
- Push-based fan-out to multiple subscribers
- No persistence (fire and forget)
- Subscribers: Lambda, SQS, HTTP, email, SMS
- Use for: event notifications, fan-out patterns

SQS (Simple Queue Service) — message queue:
- Pull-based, at-least-once delivery
- Standard (unlimited throughput, best-effort ordering)
- FIFO (exactly-once, 300 msg/s or 3000 with batching)
- 14-day retention, dead-letter queues
- Use for: decoupling, work queues, buffering

Kinesis Data Streams — real-time streaming:
- Ordered, replayable, multiple consumers
- Shard-based (1MB/s in, 2MB/s out per shard)
- 24h-365d retention
- Use for: real-time analytics, log aggregation, IoT telemetry

Pattern: SNS -> SQS fan-out for reliable event processing. Kinesis when you need ordering, replay, or real-time processing.

12. Explain AWS Lambda layers and deployment packages

Show answer Deployment package: your function code + dependencies, either as a .zip or container image.

Zip package limits:
- 50MB compressed upload
- 250MB uncompressed (including layers)
- /tmp: 512MB (configurable to 10GB)

Lambda Layers:
- Shared code/dependencies packaged separately
- Up to 5 layers per function
- Versioned and immutable
- Extracted to /opt at runtime

Use layers for:
- Shared libraries across functions (e.g., boto3 upgrade)
- Large dependencies (numpy, pandas)
- Custom runtimes

Container image alternative:
- Up to 10GB image size
- Full control over runtime environment
- Use ECR for storage
- Slower cold starts than zip

Best practice: use layers for shared deps, zip for small functions, containers for large/complex dependencies. Pin layer versions in production.

13. What is AWS CloudFormation drift detection?

Show answer Drift detection identifies resources whose actual configuration differs from the CloudFormation template definition.

# Detect drift on a stack:
aws cloudformation detect-stack-drift --stack-name mystack
aws cloudformation describe-stack-drift-detection-status --stack-drift-detection-id

Drift statuses:
- IN_SYNC: actual matches template
- MODIFIED: properties changed outside CF
- DELETED: resource removed outside CF
- NOT_CHECKED: drift detection not supported for this resource type

Common drift causes:
- Manual console changes
- CLI/SDK modifications
- Other IaC tools modifying same resources

Remediation:
1. Update template to match actual state (import drift)
2. Update stack to overwrite drift (stack update)
3. Import existing resources into the stack

Not all resource types support drift detection. Use AWS Config rules for continuous drift monitoring.

🔴 Hard (1)

1. What is the difference between DynamoDB GSI and LSI?

Show answer GSI (Global Secondary Index):
- Different partition key AND/OR sort key from base table
- Separate provisioned throughput (own RCU/WCU)
- Eventually consistent reads only
- Can be added/removed anytime
- Limit: 20 per table

LSI (Local Secondary Index):
- Same partition key, different sort key
- Shares base table throughput
- Supports strongly consistent reads
- Must be created at table creation time
- Limit: 5 per table
- 10GB partition limit (per partition key value)

Choose GSI when: querying across partition keys, need flexible access patterns.
Choose LSI when: same partition key with different sort orders, need strong consistency.

GSI is more flexible and commonly used. LSI is a constraint — prefer GSI unless you specifically need strong consistency on the alternate sort key.