Amazon ElastiCache — Redis vs Memcached Architectures
Recap: Day 26
Day 26 mapped the S3 storage-class ladder — Standard, Intelligent-Tiering, Standard-IA, One Zone-IA, and the three Glacier tiers — and the lifecycle transition rules that move objects down that ladder automatically as they age. The takeaway was that storage classes trade cost against retrieval latency and availability, that lifecycle rules automate the transitions and expirations, and that CRR and SRR replicate objects asynchronously for DR or data residency. The interesting part of that lesson was not the price list; it was the realization that durability and availability are separate axes. One Zone-IA is cheap because it gives up AZ-level durability, while Glacier Deep Archive is cheap because it gives up retrieval speed — and neither trade is reversible after the fact.
Today sits at the same lifecycle stage but addresses a different concern. S3 is durable storage for data you must keep; ElastiCache is volatile storage for data you can afford to lose, deployed specifically so that you do not have to read it from a durable store on every request. The design questions are inverted. With S3 you ask how long to keep an object and how fast you might need it back. With ElastiCache you ask how much of the working set fits in memory, what happens to a request when a node disappears, and whether the cache is allowed to be empty. That last question is the one that separates Redis from Memcached, and it is where most exam scenarios land.
Foundations You'll Need Today
Today's lesson is about caching, and it leans on a handful of ideas that are second nature to someone who has run a production system but are rarely spelled out in AWS documentation. None of them are complicated, but the rest of this page will not make sense without them, so here is the grounding.
What a Cache Is and Why Anyone Builds One
A cache is a small, fast store that holds a copy of data that also lives somewhere slower and more authoritative. The problem it solves is a mismatch between how fast your data store can answer and how often you ask it the same question. A database reading from disk might take ten milliseconds to answer a query; the same answer held in memory might take a tenth of a millisecond. If ninety percent of your requests are asking for the same small set of records, you can answer almost all of them from memory and only touch the database for the rest. That is the entire idea. The application checks the cache first, and if the value is not there — a "miss" — it reads from the database and writes the result back into the cache for next time. This pattern is called cache-aside, and it is what almost every scenario on this page is describing. The fraction of requests answered from the cache is the "hit rate," and it is the single number that tells you whether the cache is doing its job.
In-Memory vs. On-Disk Storage
Everything a computer stores lives in one of two places: memory (RAM) or disk. Memory is fast and temporary — it holds whatever is currently in use, and its contents vanish when the machine loses power. Disk is slower and persistent — it survives restarts, and it is where databases keep the data you cannot afford to lose. A cache lives in memory by definition, which is why it is fast, and that same fact is why it is fragile. When a cache node restarts, whatever it was holding is simply gone. This is not a bug to be fixed; it is the tradeoff you are accepting in exchange for the speed. The whole Redis-versus-Memcached question later on this page is really a question about how much you care that the cache's contents can disappear.
Replication, Failover, and Primary/Replica Topology
If you have one copy of something and that copy dies, you have nothing. The standard defense is to keep a second copy that stays in sync with the first, so that if the first one fails, the second can take over. The original is called the primary (or leader, or master, depending on the system) and the copy is called a replica. Keeping them in sync is replication. When the primary fails and a replica is promoted to take its place, that is a failover, and if the system does it automatically without a human intervening, it is automatic failover. The catch is that replication is usually asynchronous — the replica is told about each write a moment after the primary accepts it — so a replica is always slightly behind. That gap is called replication lag, and it matters because a failover to a replica that was behind means the writes in that gap are lost. This is exactly the mechanism behind the "users got logged out after failover" scenario later in this page.
Sharding: Splitting Data Across Multiple Machines
One machine can only hold so much data and handle so many operations per second. When you outgrow a single machine, the usual answer is to split the data across several machines and have each one own a slice of it. That is sharding, also called horizontal partitioning. The tricky part is deciding which piece of data goes where, because every client needs to be able to find any given item without asking every machine. The common approach is to run the item's key through a hash function and use the result to pick a machine — the same key always lands on the same machine, so lookups are deterministic. Redis cluster mode does exactly this, dividing the keyspace into 16,384 numbered "hash slots" and assigning ranges of slots to shards. The cost of sharding is that operations spanning multiple keys may end up on different machines, which is why the page warns that multi-key transactions are restricted in cluster mode.
TTL and Eviction: Deciding What to Throw Away
A cache has a fixed amount of memory, and eventually it fills up. Two mechanisms decide what happens next. The first is a TTL, or time-to-live: when you write a value, you can attach an expiry, and the cache will delete it automatically once that time passes. TTLs keep stale data from lingering and keep the cache from growing without bound. The second is eviction: when memory is full and a new write arrives, the cache has to remove something to make room, and the eviction policy decides what. The most common policy is LRU, least-recently-used, which throws out whatever has not been touched for the longest. The alternative is to refuse the new write instead, which is what the noeviction policy does — and as the page will explain, that is almost always the wrong choice for a cache, because it turns a capacity problem into an application error.
With that grounding — a fast but forgettable copy of your data, kept in sync by replication, split across machines by sharding, and pruned by TTLs and eviction — here is why ElastiCache exists and what problem it actually solves.
1. Why This Is on the Exam
ElastiCache appears on SAP-C02 in two distinct shapes. The first is a straightforward performance question: a relational database is saturating under read load, the workload is read-heavy with a small hot working set, and the correct answer is to put an in-memory cache in front of it rather than scale the database vertically or add read replicas. The second shape is harder and more common at the professional level — a scenario that describes a cache and then adds a constraint that eliminates one of the two engines. The constraint is almost never about raw throughput. It is about whether the cached data must survive a node failure, whether the application needs data structures beyond simple key-value, whether the cache must span regions, or whether the client library can tolerate a node disappearing from the cluster.
This maps to Domain 2, Design for New Solutions, and to a lesser extent Domain 4, Continuous Improvement for Existing Solutions. The exam expects you to know that ElastiCache is a managed deployment of Redis or Memcached, that the two engines are not interchangeable, and that the choice is driven by durability and feature requirements rather than by performance. It also expects you to recognize the cache-aside pattern in a scenario description and to know that a cache is not a system of record. A surprising number of distractors propose using ElastiCache as the primary data store for a workload that cannot tolerate data loss; that answer is always wrong, and recognizing it quickly is worth real time on the exam.
The third reason this topic is tested is that it forces a decision about failure semantics. A cache that loses its contents on failover is fine for a product catalog and catastrophic for a session store that holds authentication state. The exam will describe the workload's tolerance for cache loss in prose — "users must not be logged out during a failover," "the application can rebuild the cache from the database" — and expect you to translate that prose into an engine choice. Reading for that sentence is the skill being tested, more than any specific configuration parameter.
2. How the Two Engines Actually Work
Redis is a single-threaded, in-memory data structure server. That single-threaded execution model is not a limitation to work around; it is the reason Redis can offer atomic operations on complex types without locking. Every command executes to completion before the next one begins, so an INCR, a list push, or a sorted-set range query is atomic by construction. Redis stores values as typed structures — strings, hashes, lists, sets, sorted sets, bitmaps, HyperLogLogs, and streams — and exposes commands that operate on those structures server-side. A sorted set used as a leaderboard, for example, is updated and queried entirely inside Redis, which is why the pattern is so much cheaper than pulling the whole set into the application and sorting it there.
Memcached is a multi-threaded, in-memory key-value store with a deliberately minimal command set: get, set, add, replace, append, prepend, incr, decr, and a handful of statistics and flush operations. Values are opaque byte blobs. There is no server-side logic, no persistence, and no replication. Memcached scales by adding nodes and letting the client library shard keys across them via consistent hashing. Each node is independent; there is no cluster membership protocol, no gossip, and no concept of a primary. This makes Memcached extremely simple to operate and extremely predictable under load, because there is no cross-node coordination to go wrong.
The architectural consequence is that Redis can be made durable and highly available while Memcached cannot. Redis supports asynchronous replication to one or more replicas, optional AOF and RDB persistence to disk, automatic failover via a primary-replica topology, and cluster mode for horizontal sharding with replicas per shard. Memcached supports none of these. When a Memcached node fails, the keys it held are gone, and the client library rehashes the remaining keyspace across the surviving nodes — which means a large fraction of keys miss simultaneously and the backing database absorbs a spike. That thundering-herd behavior is the single most important operational difference between the two engines, and it is the reason a session store with a failover requirement can never be Memcached.
3. The Core Decision Boundary
Every ElastiCache scenario reduces to one question: can the application tolerate losing the entire cache contents at any moment, without a user-visible failure? If the answer is yes, Memcached is a legitimate and often cheaper choice. If the answer is no — because the cache holds session state, because a cold cache would overwhelm the database, or because the cached data is expensive to recompute — then Redis is the only option, and the remaining decisions are about how much durability and availability to buy.
The second-order question, once Redis is chosen, is whether the workload needs cluster mode. A single-shard Redis replication group with one primary and one or two replicas gives you failover and a single logical keyspace, but all writes go to one node and the dataset must fit in that node's memory. Cluster mode shards the keyspace across up to 500 shards, each with its own primary and replicas, which raises both the memory ceiling and the write throughput ceiling. The cost is that multi-key operations across hash slots are restricted, and the client must be cluster-aware. The exam tends to test this boundary with a scenario that mentions a dataset larger than a single node's memory or a write rate that one primary cannot absorb.
| Requirement in the scenario | Redis | Memcached |
|---|---|---|
| Cache may be lost without user impact | Works, but over-provisioned | Correct choice |
| Session state must survive node failure | Correct choice (replicas + failover) | Not possible |
| Sorted sets, lists, pub/sub, streams | Native | Not supported |
| Multi-threaded scaling on a single node | Single-threaded per shard | Native multi-threading |
| Cross-region replication | Global Datastore | Not supported |
| Persistence to disk | AOF / RDB snapshots | None |
| Horizontal sharding | Cluster mode, up to 500 shards | Client-side consistent hashing |
4. Configuration Modes and Their Tradeoffs
Redis on ElastiCache is deployed as a replication group, and the first decision is whether that group is cluster-mode disabled or cluster-mode enabled. With cluster mode disabled, you get exactly one shard. You can still have replicas — a primary plus up to five read replicas in the same group — and you still get automatic failover, Multi-AZ, and a single endpoint that never changes. This is the right configuration for the majority of caching workloads, because most caches are small enough to fit in one node's memory and the operational simplicity is worth a great deal. The ceiling is the node's memory and the single primary's write throughput.
Cluster mode enabled partitions the keyspace into 16,384 hash slots distributed across shards. Each shard is itself a replication group with its own primary and replicas, so you get both horizontal scale and per-shard failover. The tradeoff is that keys in different hash slots cannot participate in the same multi-key command, and transactions and Lua scripts that touch multiple keys must use hash tags to force those keys into the same slot. Client libraries must be cluster-aware. This is the configuration to reach for when the dataset exceeds a single node's memory, when write throughput must exceed what one primary can absorb, or when the scenario explicitly mentions hundreds of gigabytes of cached data.
Memcached has no equivalent topology decisions because it has no topology. You create a cluster of nodes, and the client library distributes keys across them. There is no primary, no replica, no failover, and no endpoint that survives a node replacement. Auto Discovery lets clients learn about added or removed nodes without a redeploy, which softens the operational pain of scaling, but it does not change the failure semantics. When a node is lost, its keys are lost, and the client rehashes. The configuration knobs that matter are node count and node size, and the only real design decision is how much headroom to leave so that a single node loss does not push the surviving nodes past their memory limit.
For Redis, the remaining knobs are persistence and eviction. Persistence is optional and off by default in the sense that a cache is expected to be rebuildable; enabling AOF or RDB snapshots costs write throughput and storage but shortens the cold-start window after a full cluster replacement. Eviction policy determines what happens when memory fills: volatile-lru and allkeys-lru are the common choices, and the difference is whether keys without a TTL are eligible for eviction. Choosing noeviction on a cache is almost always a mistake, because writes begin failing instead of the cache gracefully shedding cold data.
5. Sizing, Limits and Quotas
Sizing an ElastiCache deployment starts with the working set, not the total dataset. The question is how much data must be resident in memory to keep the hit rate above whatever threshold the application needs, and that number is usually far smaller than the total data volume. A cache holding the hot 10% of a catalog can serve the overwhelming majority of reads. Once the working set is known, add headroom for the eviction policy to operate — running a cache at 95% memory utilization means constant eviction churn and a hit rate that degrades unpredictably.
The limits that matter for exam reasoning are the topology ceilings. A Redis replication group with cluster mode disabled supports one shard with up to five read replicas. With cluster mode enabled, a cluster supports up to 500 shards, and each shard can have its own replicas. Memcached clusters support up to 20 nodes per cluster, which is a hard ceiling that occasionally appears in scenarios describing very large caches. Both engines support node types ranging from small burstable instances up to memory-optimized families, and the largest node types are what you reach for when a single shard must hold a very large working set.
| Dimension | Redis (cluster mode disabled) | Redis (cluster mode enabled) | Memcached |
|---|---|---|---|
| Shards per cluster | 1 | Up to 500 | N/A (independent nodes) |
| Replicas per shard | Up to 5 | Up to 5 | None |
| Nodes per cluster | Up to 6 | Up to 500 shards × replicas | Up to 20 |
| Automatic failover | Yes (Multi-AZ) | Yes, per shard | No |
| Cross-region replication | Global Datastore | Global Datastore | No |
| Persistence | AOF / RDB | AOF / RDB | None |
Two quotas are worth remembering because they shape architecture rather than just cost. First, scaling a cluster-mode-disabled Redis group vertically requires a node type change, which is an in-place operation with a brief failover; scaling a cluster-mode-enabled group horizontally is an online resharding operation that moves hash slots between shards. Second, Memcached's 20-node ceiling means that a workload needing more than 20 nodes of cache capacity must either use larger nodes or move to Redis cluster mode. Scenarios that describe a cache growing past a fixed node count are usually pointing at that ceiling.
6. Failure Modes and What They Look Like in Production
The most common production failure is not a node dying; it is the cache being cold. A deployment, a node replacement, or a failover that lands on a replica with stale or empty data produces a sudden drop in hit rate, and the backing database absorbs the full read load. The symptom is a database CPU spike correlated with a cache event, and the first diagnostic move is to check the cache hit rate metric alongside database read throughput on the same time axis. If they move together, the cache is the cause and the database is the victim.
The second failure mode is memory pressure leading to eviction churn. When a cache runs near its memory limit, the eviction policy starts removing keys faster than the workload's access pattern can tolerate, and the hit rate collapses even though the cache is technically healthy. The metrics to watch are Evictions, CurrItems, and BytesUsedForCache. A rising eviction count with a flat item count means the working set no longer fits, and the fix is either a larger node or a smaller TTL, not more nodes.
The third failure mode is specific to Memcached and is the one exam scenarios are built around. When a Memcached node is lost, the client library rehashes the keyspace across the surviving nodes. Every key that was on the lost node now misses, and every key that moved to a different node also misses because the client is looking in the wrong place. The result is a simultaneous miss storm across a large fraction of the keyspace, which is far worse than the proportional loss you might expect. Redis avoids this because a replica takes over the same hash slots, so the keyspace mapping does not change. This is the concrete mechanism behind the abstract statement that Redis survives node failure and Memcached does not.
The fourth failure mode is connection exhaustion. Both engines have a maximum number of client connections per node, and a fleet of application instances that each open a connection pool can exhaust it during a scale-out event. The symptom is connection errors from the client library while the cache itself reports healthy CPU and memory. The fix is connection pooling on the client side, and the metric to alarm on is CurrConnections approaching the node's limit.
7. The Operational and SRE Angle
A cache is a latency optimization with a failure mode that is worse than not having it, because the backing store is sized for the cached path and not for the uncached one. That asymmetry is the core SRE concern. If the database can absorb full uncached load, the cache is a pure optimization and its failure is a performance incident. If it cannot, the cache is a load-bearing component and its failure is an availability incident. Deciding which of those two you have built is the first thing to establish, because it determines whether the cache needs Multi-AZ, replicas, and alarms, or whether a single node is genuinely sufficient.
The metrics worth alarming on are hit rate, eviction rate, memory utilization, current connections, and replication lag for Redis. Hit rate is the leading indicator: a sustained drop below the application's tolerance threshold is the earliest signal that something has changed, whether that is a cold cache, a working-set growth, or a bad deployment that changed key patterns. Eviction rate is the capacity signal. Replication lag matters because a failover to a replica that is far behind means losing recent writes, which for a session store means logging users out.
The runbook shape follows from the failure modes. For a cold cache, the runbook is to confirm the cache event, then decide whether to warm the cache deliberately or let it warm organically while protecting the database — often by temporarily reducing application concurrency or enabling a request-level circuit breaker. For memory pressure, the runbook is to identify the largest key families and either raise the node size or shorten TTLs. For a Memcached node loss, the runbook is to expect a miss storm and to have a database-side protection mechanism ready, because there is no cache-side fix. For Redis failover, the runbook is to verify the new primary, confirm replication is re-established, and check whether the application reconnected cleanly.
8. Edge Cases and Exam Gotchas
The single most reliable gotcha is the assumption that ElastiCache is a database. It is not. It has no query language, no transactions across arbitrary keys, no schema, and no durability guarantee unless you explicitly enable persistence — and even then, persistence is a recovery aid, not a system of record. Any scenario that describes data that must not be lost and proposes ElastiCache as the store is wrong, regardless of how the rest of the architecture is described.
The second gotcha is the belief that Redis is always the right answer. Redis is the right answer whenever durability, replication, or data structures matter, but a scenario describing a simple, horizontally scaled, stateless cache where a node loss is acceptable is describing Memcached, and choosing Redis there is over-engineering. The exam does test this direction, not just the Redis-is-better direction.
The third gotcha is cluster mode. Enabling cluster mode changes the client contract: multi-key operations must use hash tags, and the client library must be cluster-aware. A scenario that describes an application using multi-key transactions and a cache that fits comfortably in one node is describing cluster mode disabled, even if the dataset is large in absolute terms. The trigger for cluster mode is a working set that exceeds one node's memory or a write rate that one primary cannot absorb, not total data volume.
The fourth gotcha is Global Datastore. It provides cross-region replication for Redis, and it is the answer when a scenario requires a cache to be available in a second region with low-latency local reads. It is not a DR mechanism for the cache in the sense that it guarantees zero data loss; it is a replication feature with asynchronous propagation. The fifth gotcha is that Memcached's Auto Discovery solves node discovery, not node failure. It lets clients learn about topology changes without a redeploy, but it does not preserve data when a node is lost, and scenarios sometimes conflate the two.
9. ElastiCache vs. the Services It Gets Confused With
ElastiCache is most often confused with DynamoDB Accelerator, with self-managed Redis on EC2, and with the database it is caching. DAX is the closest analogue and the most common distractor: it is an in-memory cache, but it is purpose-built for DynamoDB and speaks the DynamoDB API, so it only helps when the backing store is DynamoDB. If the scenario's backing store is Aurora or RDS, DAX is not applicable and ElastiCache is the answer. If the backing store is DynamoDB and the workload is read-heavy on hot items, DAX is the answer and ElastiCache would require application-level cache logic that DAX provides transparently.
Self-managed Redis on EC2 is the other common distractor. It offers the same engine and therefore the same feature set, but it moves patching, failover orchestration, backup, and monitoring onto the team. The exam generally prefers the managed service unless the scenario explicitly requires a Redis version, module, or configuration that ElastiCache does not support. That is a real constraint — ElastiCache does not support arbitrary Redis modules — but it is stated explicitly when it is the deciding factor.
| Option | Backing store | Pick it when… |
|---|---|---|
| ElastiCache for Redis | Any (Aurora, RDS, DynamoDB, custom) | You need replication, failover, data structures, pub/sub, or cross-region replication |
| ElastiCache for Memcached | Any | The cache is simple key-value, horizontally scaled, and node loss is tolerable |
| DAX | DynamoDB only | The backing store is DynamoDB and you want transparent, API-compatible caching |
| Self-managed Redis on EC2 | Any | You need a Redis module or version ElastiCache does not offer |
| Read replicas (RDS/Aurora) | Relational | The read load is broad rather than hot-key, and staleness is acceptable |
The rule to carry into the exam is short. Choose Redis when the cache must survive failure, when the application needs server-side data structures, or when the cache must span regions. Choose Memcached when the cache is a pure, disposable, horizontally scaled key-value layer and the application can rebuild it. Choose DAX when the backing store is DynamoDB. Choose read replicas when the problem is read volume rather than hot-key latency, because a cache and a read replica solve different problems even though both reduce database load.
Hands-On Lab: Session Store Failover, Redis vs. Memcached (60 min)
The goal of this lab is to build the same session-store workload twice — once on Redis and once on Memcached — and then break each one deliberately, so that the difference in failure behavior is something you have observed rather than something you have read. Work in a sandbox account. Everything here is small enough to run inside the free-tier-adjacent node types, and the whole lab should cost a few dollars if you tear it down the same day.
1. Establish the baseline workload. Deploy a small application — a container on ECS Fargate or a t3.small EC2 instance is sufficient — that exposes two endpoints: one that writes a session token to the cache with a 30-minute TTL, and one that reads it back and reports whether the read was a hit or a miss. Point it at a backing store that can also persist sessions, so you have a fallback path to compare against. Instrument the application to emit a hit/miss counter to CloudWatch.
2. Build the Redis session store. Create an ElastiCache for Redis replication group with cluster mode disabled, one primary and one replica, Multi-AZ enabled, and automatic failover turned on. Note the primary endpoint. Configure the application to use it, then drive a steady read/write load with a simple loop or a load-testing tool for several minutes so that the cache is warm and the hit rate is stable.
3. Force a Redis failover. Trigger a failover on the replication group — either through the console's failover action or by rebooting the primary with failover. Watch the CloudWatch metrics for the replication group: CurrConnections, CacheHits, CacheMisses, and the failover event in the ElastiCache event log. Record how long the application saw errors, whether the session tokens written before the failover were still readable afterward, and whether the client reconnected without a restart. This is the behavior you are buying when you choose Redis.
4. Build the Memcached session store. Create an ElastiCache for Memcached cluster with three nodes and Auto Discovery enabled. Point the application at the cluster's discovery endpoint and repeat the same warm-up load. Confirm the hit rate stabilizes and that keys are distributed across all three nodes by checking the per-node CurrItems metric.
5. Kill a Memcached node. Terminate one of the three nodes and watch what happens. The client library will rehash the keyspace across the two survivors. Record the hit-rate drop, the corresponding spike in backing-store reads, and how long the cache takes to re-warm. Compare that miss storm against the Redis failover you measured in step 3. The difference is the entire argument for choosing one engine over the other.
6. Test the eviction boundary. On the Redis group, set the eviction policy to allkeys-lru and drive writes until BytesUsedForCache approaches the node's memory limit. Watch Evictions climb and the hit rate degrade. Then change the policy to noeviction and observe that writes begin failing instead. This is the tradeoff between shedding cold data and rejecting new writes, and it is worth seeing once.
7. Tear down and write up. Delete both clusters, the application, and any load-testing resources. Write a short comparison table with the numbers you measured: failover duration, hit-rate drop, recovery time, and error count for each engine. Keep it — the numbers are more persuasive than the prose when you are justifying an engine choice in a design review.
Scenario Question Drills (15 questions)
Q1. A session store must survive a node failure without losing data and support automatic failover. Which caching engine and configuration fits?
Q2. A product catalog cache holds the hot 5% of a 400 GB catalog. The application can rebuild the cache from Aurora in under a minute and users see no error during a rebuild. Cost is the primary constraint. Which engine should you choose?
Q3. A gaming leaderboard must be updated and queried atomically, returning the top 100 players by score in a single call. Which ElastiCache configuration fits?
Q4. A cache must hold 900 GB of working set and absorb 400,000 writes per second. A single Redis node type tops out well below both. What should you configure?
Q5. A Memcached cluster loses one of its three nodes. What is the expected application-visible behavior?
Q6. A read-heavy application uses DynamoDB as its backing store and needs microsecond read latency on hot items with no application-level cache logic. What should you deploy?
Q7. A Redis cache is running at 96% memory utilization. Hit rate has dropped sharply over the past hour and Evictions is climbing steadily. What is the correct first response?
noeviction would make writes fail, and persistence does not address capacity.Q8. An application uses Redis multi-key transactions across several related keys. The dataset fits comfortably in a single node's memory. Which configuration should you choose?
Q9. A globally distributed application needs its cache readable with low latency in a second region, with writes in the primary region propagating to the secondary. Which feature provides this?
Q10. A team wants to use a Redis module that ElastiCache does not support. What is the appropriate architecture?
Q11. A cache is sized so that a single node loss would push the surviving Memcached nodes past their memory limit. What is the correct design change?
Q12. A cache is deployed with noeviction and begins returning write errors as memory fills. What is the underlying problem?
noeviction, Redis returns errors on writes once memory is full rather than evicting cold data. For a cache, an LRU policy that sheds cold keys is nearly always the better behavior.Q13. After a Redis failover, users report being logged out even though the failover completed successfully. What is the most likely cause?
Q14. A workload needs a cache that supports publish/subscribe messaging between application instances in addition to key-value storage. Which engine fits?
Q15. A scenario describes a cache that must not lose data, must be queryable by the application with complex filters, and must serve as the authoritative record for user profiles. What is wrong with proposing ElastiCache?
Peek into Tomorrow
Today's decision boundary was narrow and specific: given a caching requirement, choose Redis or Memcached. Tomorrow widens the frame considerably. The question becomes how to choose among an entire portfolio of data services — relational, key-value, document, cache, object — when a scenario describes latency, consistency, and scale requirements without naming a service. That is a harder problem, because the constraints in those scenarios are often in tension. A workload that needs single-digit-millisecond latency at unpredictable scale points at DynamoDB, but a workload that needs multi-region active writes points at DynamoDB Global Tables specifically, and a workload that needs relational semantics with cross-region reads points at Aurora Global Database instead. The two global replication features sound similar and behave very differently: one is multi-active with last-writer-wins conflict resolution, the other is single-writer with a promotable secondary.
The unresolved question today leaves open is where the cache fits in that matrix. A cache is not a data service in the same sense as the others — it is a layer in front of one — and the synthesis has to place it correctly rather than treating it as a peer. Tomorrow's scenario drills are built to expose exactly that confusion, along with the consistency-model tradeoffs that separate the relational and key-value families. If today's material felt like a narrow engine comparison, tomorrow is where it becomes a selection framework.
Sources
- Amazon ElastiCache — What Is ElastiCache?
- Choosing Between Memcached and Redis
- ElastiCache for Redis — Replication Groups and Failover
- ElastiCache for Redis — Cluster Mode
- ElastiCache for Redis — Global Datastore
- ElastiCache — Cache Nodes and Node Types
- ElastiCache — Monitoring with CloudWatch Metrics
- ElastiCache — Best Practices
- Amazon DynamoDB Accelerator (DAX)
- ElastiCache — Caching Strategies