AWS Snow Family — Snowball Edge, Snowcone & Snowmobile
Recap: Where We Left Off
Day 48 covered Storage Gateway, and the whole lesson turned on a single question: what protocol does the on-prem workload actually speak? A File Gateway presents S3 as an NFS/SMB share with local caching, a Volume Gateway exposes iSCSI block storage backed by S3 in either cached or stored mode, and a Tape Gateway presents a virtual tape library so existing backup software keeps working unchanged. The takeaway was that the three flavors are not interchangeable — you pick based on whether the workload speaks file, block, or tape, and the gateway is a long-lived bridge that stays in the path.
Today sits at the same lifecycle stage but addresses a different concern. Storage Gateway assumes a working, persistent network path between on-prem and AWS; it is a bridge you keep. The Snow Family is the opposite posture — it exists precisely for the case where the network path is the bottleneck, absent, or so slow that moving the bytes physically is faster than moving them over the wire. Same migration phase, same "get data from here to there" problem, but the constraint being solved is bandwidth rather than protocol translation.
Foundations You'll Need Today
Bandwidth math: why "how big" is only half the question
Network links are measured in bits per second, while data volumes are measured in bytes, and the two are off by a factor of eight. A link rated at 100 Mbps moves 100 megabits per second, which is about 12.5 megabytes per second, which works out to roughly 1 terabyte per day if the link runs flat out with no other traffic. That last clause is the important one: real links are shared, and real transfers carry protocol overhead, so the practical number is usually 60-80% of the theoretical one. The reason this matters today is that the entire Snow Family decision reduces to a single division — total data volume divided by the rate the link can actually sustain — compared against the deadline the business will tolerate. If that division produces a number larger than the deadline, the network is not a viable path and you need a different mechanism.
Amazon S3: the destination everything is moving toward
Amazon Simple Storage Service (S3) is AWS's object storage service — a place to put files where you never think about disks, servers, or file systems. You create a bucket (a named container), and you put objects (files plus their metadata) into it. There is effectively no capacity limit on a bucket, you pay per gigabyte stored and per request made, and the data is replicated across multiple facilities within a Region automatically. Almost every migration in this phase of the curriculum ends with data landing in S3, because S3 is the cheapest durable place to park bulk data and because most other AWS services can read from it directly. When this lesson says a Snowball's contents are "ingested into S3," it means AWS takes the returned device, copies its contents into a bucket you specified when you created the job, and then wipes the device.
Encryption at rest, and who holds the keys
"Encryption at rest" means the data is scrambled on the storage medium itself, so that someone who physically steals the disk sees meaningless bytes rather than your files. The subtlety is key custody: if the decryption key is stored on the same device as the encrypted data, a thief who takes the device has both halves of the puzzle and the encryption buys you very little. The stronger model — the one the Snow Family uses — keeps the keys entirely inside AWS, never on the device. The device can encrypt as it writes, but it cannot decrypt what it has written, because it never had the key. This is why a lost Snowball is an operational inconvenience rather than a data breach, and it is why the tamper-detection hardware matters: if someone opens the case, the device destroys what little key material it holds and the data becomes permanently unreadable.
EC2 instances and Lambda functions: the compute that can run on the device
Amazon EC2 is AWS's virtual server service — you rent a machine in an AWS data center, choose its size and operating system, and run whatever you want on it. AWS Lambda is the serverless counterpart: you upload a function (a small piece of code), and AWS runs it on demand without you managing any server at all. Both normally run inside an AWS Region, which is the point of the next concept: the Snowball Edge Compute Optimized variants can run EC2 instances and Lambda functions on the device itself, out at a remote site with no Region connectivity. That is what turns a data mover into an edge-computing platform, and it is why a scenario mentioning local processing or inference points at the same product family as a scenario about bulk transfer.
With that grounding, here's why the Snow Family exists and what problem it actually solves.
1. Why the Snow Family Is on the Exam
The Snow Family exists because of a hard physical constraint that no amount of AWS service design can route around: the speed of light through fiber, multiplied by the cost of the bandwidth you can actually buy. A 100 Mbps uplink moves roughly 1 TB per day in theory, and considerably less in practice once you account for TCP overhead, competing traffic, and the fact that most enterprise WAN links are shared. At that rate, a 500 TB dataset takes months. The same dataset shipped on physical devices takes days, including transit time. The exam tests whether you recognize when that crossover point has been reached and which device class fits the volume.
This maps most directly to SAP-C02 Domain 3, Migration and Modernization, and specifically to the "determine the right migration strategy" task. The exam does not ask you to configure a Snowball in detail. It asks you to read a scenario containing a data volume, a network constraint, and a timeline, and then choose between online transfer (DataSync, Direct Connect, S3 Transfer Acceleration), offline transfer (Snow Family), or a hybrid of both. The distractor pattern is consistent: options that would work in principle but violate the stated constraint, such as "provision a Direct Connect" when the scenario says the facility has no fiber path, or "use DataSync" when the scenario says the link is saturated for months.
There is a second, less obvious exam angle. The Snow Family devices are not just data movers — the Compute Optimized variants run EC2 instances and Lambda functions locally, which makes them an edge-computing answer as well as a migration answer. A scenario describing a ship, a mine site, or a factory floor with intermittent connectivity and a need to process data before shipping it back is testing the same service from the edge-compute direction. Recognizing that both angles point at the same product family is worth a question or two.
Finally, the exam expects you to know the rough capacity tiers well enough to size an engagement. You do not need exact SKU numbers, but you do need to know that Snowcone is measured in single-digit terabytes, Snowball Edge in tens of terabytes, and Snowmobile in petabytes. A scenario that says "2 PB" and offers Snowball Edge as an option is testing whether you know that Snowmobile exists for that scale, or whether you know that a fleet of Snowball Edge devices is a legitimate alternative answer when the scenario emphasizes cost or timeline over a single shipment.
2. How the Devices Actually Work
Every Snow Family device is, at its core, a ruggedized server with local storage, a CPU, and a set of network interfaces, shipped to you in a case that survives freight handling. The workflow is deliberately low-tech on the AWS side: you order a job through the console or API, AWS ships the device, you connect it to your local network, copy data onto it using the Snowball client or the OpsHub application, ship it back, and AWS ingests the contents into S3. The device is wiped after ingestion using NIST 800-88 standards, and you receive a job completion report.
The security model is worth understanding because it shows up in exam scenarios about regulated data. Data on the device is encrypted at rest with 256-bit encryption, and the keys never leave AWS — they are not stored on the device itself. The device is tamper-resistant and tamper-evident, with a Trusted Platform Module that validates the device's integrity at boot. If the device is opened or the TPM detects tampering, the keys are destroyed and the data is unrecoverable. This is why the Snow Family is acceptable for data that would otherwise be blocked from leaving a facility: the encryption and key custody model is stronger than most on-prem tape workflows.
For the Compute Optimized variants, the device is more than a bucket. It runs the AWS IoT Greengrass service and can execute EC2 instances and Lambda functions locally, which means you can run inference, filtering, or aggregation on the device before shipping it back. This matters for the edge scenarios: a ship with a satellite link that costs per megabyte does not want to ship raw sensor data, it wants to run a model locally and ship the results. The same device that moves bulk data can also be the compute node that decides what is worth moving.
OpsHub is the piece that makes this usable without a laptop full of scripts. It is a graphical application that runs on a local workstation, connects to the device over the local network, and provides a file-transfer interface, job monitoring, and the ability to launch compute instances on the device. For a data center operator who is not a developer, OpsHub is the difference between a workable migration and a support ticket. The exam rarely asks about OpsHub directly, but it appears as a distractor option and you should know it is the management UI, not a transfer service.
3. The Core Decision Boundary: Online vs. Offline
Almost every Snow Family exam question reduces to one fork: should this data move over the network or physically? The answer is not about volume alone — it is about the ratio of volume to available bandwidth, measured against the timeline the business will tolerate. A 10 TB dataset over a 10 Gbps link is an afternoon's work and should never touch a Snowball. The same 10 TB over a 10 Mbps link is weeks of transfer and is a clear offline case. The exam gives you both numbers and expects you to do the division.
The second-order consideration is whether the transfer is one-time or ongoing. Snow devices are inherently one-shot: you ship them, they come back, the job ends. If the scenario describes a continuous stream of new data — daily logs, ongoing replication, a live feed — then offline transfer solves only the historical backlog, and you still need a network path for the delta. This is the hybrid pattern that shows up repeatedly: bulk historical data on Snowball, ongoing changes over Direct Connect or DataSync. Recognizing that the question is asking about the backlog specifically, not the whole pipeline, is often the difference between the right and wrong answer.
A third consideration is whether the data can leave the facility at all. Some regulated environments prohibit shipping storage media off-site regardless of encryption. In those cases the Snow Family is disqualified and the answer has to be a network-based option, even if it is slow. The exam occasionally plants this constraint in a single sentence and expects you to notice that it overrides the bandwidth math entirely.
| Scenario signal | Right answer | Why |
|---|---|---|
| Large volume, adequate bandwidth, ongoing | DataSync or Direct Connect | Network path exists and the transfer is continuous |
| Large volume, inadequate bandwidth, one-time | Snowball Edge | Physical transfer beats the wire |
| Petabyte-scale, single engagement | Snowmobile or a Snowball Edge fleet | Device capacity and shipment logistics |
| Small volume, rugged or remote site | Snowcone | Portable, low-capacity, edge-capable |
| Bulk historical plus ongoing delta | Snowball for bulk, DataSync/DMS for delta | Offline solves the backlog, network solves the stream |
| Media cannot leave the facility | Network transfer only | Regulatory constraint overrides bandwidth math |
4. Device Classes and What Each Costs You
The Snow Family is not one product with three sizes; it is three products with different design centers, and the tradeoffs between them are about more than capacity. Snowcone is the smallest and most portable, designed for edge locations where you need a small amount of storage and possibly some compute in a device you can carry. It is measured in single-digit terabytes and is the right answer when the scenario emphasizes portability, harsh environments, or a site too small to justify a full Snowball. The tradeoff is capacity: you will not move a data center with a Snowcone.
Snowball Edge comes in two flavors that matter for the exam. Storage Optimized is the bulk data mover, with usable capacity in the tens of terabytes per device and a design that prioritizes throughput and storage density. Compute Optimized trades some storage for more CPU and an optional GPU, and is the right answer when the scenario needs local processing — video analysis, ML inference, signal processing — at the edge. Both variants support clustering, which lets you treat multiple devices as a single logical storage pool and increases both capacity and throughput.
Snowmobile is the outlier: a shipping container with networking and storage inside, delivered by truck, capable of moving up to 100 PB in a single engagement. It is not a device you order casually. The exam treats it as the answer for exabyte-scale or very-large-petabyte migrations where even a fleet of Snowballs would be operationally absurd. If a scenario says "100 PB" and offers Snowmobile, that is almost certainly the intended answer; if it says "2 PB," Snowmobile is technically capable but a Snowball Edge fleet is usually the more cost-effective framing.
The cost model reinforces these distinctions. You pay for the device, for shipping both ways, for the days you hold it, and for the data transfer out of S3 if you later move it. There is no per-gigabyte transfer charge for the ingest itself, which is a meaningful difference from network transfer. For very large datasets, the offline path is often cheaper as well as faster, and the exam sometimes frames the question purely on cost rather than time.
5. Sizing, Limits, and the Bandwidth Math
The single most useful number to internalize is the throughput of a typical WAN link expressed in terabytes per day. A 100 Mbps link moves roughly 1 TB per day under ideal conditions; a 1 Gbps link moves roughly 10 TB per day; a 10 Gbps link moves roughly 100 TB per day. These are order-of-magnitude figures, not guarantees, and real-world throughput is typically 60-80% of theoretical once you account for protocol overhead and competing traffic. The exam expects you to do this arithmetic quickly and compare it against the timeline in the scenario.
Against those numbers, the device capacities define the crossover. A Snowball Edge Storage Optimized device holds tens of terabytes, so a 500 TB dataset needs roughly a dozen devices or a clustered set. A Snowcone holds single-digit terabytes, so it is a fit for a few terabytes at a remote site, not for a data center evacuation. Snowmobile holds up to 100 PB, which is the answer for the scenarios that would otherwise require hundreds of Snowballs.
| Device | Rough usable capacity | Compute | Typical use |
|---|---|---|---|
| Snowcone | Single-digit TB | Limited, edge-capable | Remote or rugged sites, small transfers |
| Snowball Edge Storage Optimized | Tens of TB per device | Moderate | Bulk data migration, clustering supported |
| Snowball Edge Compute Optimized | Tens of TB, less than Storage Optimized | High, optional GPU | Edge processing, ML inference, video analysis |
| Snowmobile | Up to 100 PB | N/A | Exabyte-scale single engagements |
Two operational limits matter beyond raw capacity. First, the device is held for a limited window, and you pay for the days you keep it; a migration that stalls because the local team was not ready costs money and delays the next job. Second, the ingest into S3 is not instantaneous — AWS processes the returned device and the data appears in your bucket over a period of hours to days depending on volume. Neither limit usually changes the exam answer, but both appear in scenario details that distinguish a well-planned migration from a naive one.
6. Failure Modes and What They Look Like in Production
The most common failure in a Snow migration is not a hardware fault — it is a planning fault. The device arrives, the local team discovers that the data is spread across a dozen file systems with inconsistent permissions, and the copy takes three times longer than estimated. The symptom is a job that sits in progress well past its expected completion, and the first diagnostic move is to check the Snowball client logs for throughput and error rates rather than assuming the device is faulty. In practice, the bottleneck is almost always the source storage or the local network, not the Snowball.
The second failure mode is a mismatch between what was copied and what was expected. The Snowball client copies files, not application state, so a database that was copied while running will produce an inconsistent snapshot. The symptom is a target that looks complete but fails integrity checks or application-level validation. The fix is procedural: quiesce the source, or use a database-aware tool for the database portion and reserve the Snowball for file data. This is the same trap that catches people with raw EBS snapshots of running databases, and it is worth flagging explicitly in a runbook.
The third failure mode is a lost or damaged device in transit. This is rare, and the encryption model means the data is not exposed, but the job still has to be restarted. The symptom is a job that never reaches the "received" state, and the first move is to open a support case with the tracking information. The exam does not test this directly, but it is the reason the encryption and key-custody model matters: a lost device is an operational inconvenience, not a breach.
Finally, there is the failure mode of choosing the wrong tool. A team that ships a Snowball for a 5 TB transfer over a healthy 1 Gbps link has spent weeks of logistics on an afternoon's work. The symptom is a migration that is slower and more expensive than the network path would have been, and the diagnostic is simply the bandwidth math from the previous section. Recognizing this in a scenario — where the numbers clearly favor the network — is a common exam trap in the other direction.
7. The Operational and SRE Angle
From an SRE perspective, a Snow migration is a batch job with a long, mostly opaque middle. You cannot watch the bytes move in real time the way you can watch a DataSync task, because the device is physically in transit for days. What you can monitor is the job state in the AWS console or API, the Snowball client's local logs during the copy phase, and the S3 ingest progress after the device is returned. The runbook should define what "healthy" looks like at each phase and who is responsible for checking it.
The SLO implications are unusual. A Snow migration is not a service with an availability target; it is a project with a deadline. The relevant metric is time-to-completion against the business timeline, and the relevant alarm is a job that has not advanced state within its expected window. If the migration is feeding a cutover, the cutover date is the SLO, and every day of slippage is a direct business impact. This is why the planning phase — inventorying the data, estimating throughput, ordering enough devices — matters more than the execution phase.
The runbook shape follows the phases. Before the device arrives: confirm the source data is quiesced or that a database-aware tool is handling the live portion, confirm the local network can sustain the copy rate, and confirm someone is on-site to receive and connect the device. During the copy: monitor throughput and error rates, and escalate if the copy is falling behind the estimate. After the copy: verify the checksums, ship the device, and track the ingest into S3. After ingest: validate the target data against the source before decommissioning anything.
One operational detail worth building into the runbook is the handling of the encryption keys and the job manifest. The keys are managed by AWS, but the job manifest — what was copied, when, and to which bucket — is your audit trail. For regulated data, that manifest is often the artifact an auditor asks for, and reconstructing it after the fact is painful. Capture it at job creation, not at job completion.
8. Edge Cases and Exam Gotchas
The first gotcha is conflating the Snow Family with Storage Gateway. They solve opposite problems. Storage Gateway is a persistent bridge for ongoing access; Snow is a one-time physical transfer. A scenario that describes continuous access to on-prem data from AWS is a Storage Gateway question, and a scenario that describes moving a large dataset once is a Snow question. The exam mixes these deliberately, and the distinguishing word is usually "ongoing" versus "migrate."
The second gotcha is forgetting that Snow devices can run compute. A scenario about a remote site that needs to process data locally before sending results back is not just a storage question — it is a Compute Optimized Snowball Edge question, or possibly a Snowcone question if the volume is small. The presence of "process," "analyze," or "infer" in the scenario is the signal.
The third gotcha is the regulatory constraint. If the scenario says data cannot leave the facility, no Snow device is acceptable regardless of encryption, and the answer must be a network-based option. This constraint is often buried in a subordinate clause, and missing it is a common way to lose a question that otherwise looks straightforward.
The fourth gotcha is the hybrid pattern. A scenario that describes a large historical dataset plus ongoing changes is not asking you to choose between Snow and DataSync — it is asking you to combine them. The correct answer names both: Snow for the bulk, a network tool for the delta. Answers that pick only one are usually wrong.
The fifth gotcha is Snowmobile sizing. Snowmobile is the answer for exabyte-scale or very-large-petabyte single engagements, not for every large migration. A 2 PB migration can be done with Snowmobile, but a Snowball Edge fleet is often the more cost-effective answer, and the exam sometimes expects you to recognize that. Read the scenario for cost or timeline emphasis before defaulting to the largest device.
9. Snow vs. the Services It Gets Confused With
The Snow Family sits in a crowded space of data-movement services, and the exam expects you to pick the right one from a scenario. The distinguishing questions are always the same: is the transfer one-time or ongoing, is the network path adequate, and does the data need to be processed before it moves? Answer those three and the service choice usually falls out.
| Service | Transfer type | Network dependency | Pick it when… |
|---|---|---|---|
| Snowball Edge | One-time, physical | None for the bulk | Large volume, inadequate bandwidth, one-shot |
| Snowcone | One-time, physical | None for the bulk | Small volume, remote or rugged site |
| Snowmobile | One-time, physical | None for the bulk | Exabyte-scale single engagement |
| DataSync | Ongoing, online | Requires adequate link | Continuous or scheduled file/object sync |
| Storage Gateway | Ongoing, online | Requires adequate link | Persistent on-prem access to AWS storage |
| Direct Connect | Ongoing, online | Requires dedicated circuit | Sustained high-bandwidth hybrid connectivity |
| S3 Transfer Acceleration | Ongoing, online | Requires adequate link | Long-distance internet uploads to S3 |
The rule of thumb: if the scenario says "migrate" and gives a large volume with a constrained link, it is Snow. If it says "sync," "replicate," or "ongoing," it is DataSync or Storage Gateway. If it says "dedicated connection" or "consistent bandwidth," it is Direct Connect. If it says "accelerate uploads over the internet," it is Transfer Acceleration. The Snow Family is the only one of these that does not require a working network path, and that is the property the exam is testing.
Hands-On Lab: Sizing a Snowball Migration
Scenario. A media company has 500 TB of archived video in an on-prem data center. The facility has a 100 Mbps internet uplink shared with production traffic. The business wants the archive in S3 within 60 days, and the data cannot be compressed further. Your job is to decide whether to transfer over the network or use Snowball Edge devices, and to size the engagement.
Step 1 — Establish the network baseline. Calculate the theoretical transfer time for 500 TB over 100 Mbps. At 100 Mbps, the link moves roughly 1 TB per day under ideal conditions, so 500 TB is roughly 500 days. Apply a realistic efficiency factor of 60-80% to account for protocol overhead and the fact that the link is shared with production traffic. The practical estimate is 600-800 days. Write this number down; it is the baseline every other option is compared against.
Step 2 — Compare against the deadline. The business wants the data in S3 within 60 days. The network path misses that by more than an order of magnitude. This is the crossover point: the network is not a viable option, and the question becomes which Snow device and how many.
Step 3 — Size the device fleet. Assume a Snowball Edge Storage Optimized device holds roughly 80 TB usable. Divide 500 TB by 80 TB to get approximately 7 devices. Add one for margin, since real-world usable capacity is lower than the nominal figure and you do not want to run a device to 100% full. A fleet of 8 devices is a reasonable plan. Note that the devices can be clustered, which simplifies the copy and increases aggregate throughput.
Step 4 — Estimate the copy time. The copy rate is limited by the local network and the source storage, not the Snowball. If the local network can sustain 1 Gbps to the device, each device takes roughly 80 TB / (1 Gbps) ≈ 7-8 days to fill. Running devices in parallel shortens the wall-clock time; running them serially does not. Plan for the copy phase to take 2-3 weeks with parallel copies, plus shipping time in both directions.
Step 5 — Plan the cutover. The archive is not live data, so there is no delta to sync — this is a clean one-time migration. Confirm that the source data is quiesced (no new writes during the copy), capture the job manifest for audit, and define the validation step: after ingest, compare object counts and checksums between source and S3 before decommissioning the on-prem archive.
Step 6 — Document the decision. Write a one-page decision record: the network baseline, the deadline, the device count, the copy plan, and the validation criteria. This is the artifact that survives the migration and answers the "why did we do it this way" question six months later. It is also the artifact an auditor asks for if the data is regulated.
Scenario Question Drills
Q1. A company must migrate 2 PB of data from a facility with no viable network uplink for bulk transfer. What is the appropriate approach?
Q2. A remote research station generates 4 TB of sensor data per month and has only a satellite link. The data must be processed locally before transmission to reduce bandwidth cost. Which device fits best?
Q3. A team needs to move 500 TB of archived data to S3. The facility has a 100 Mbps uplink shared with production traffic, and the business wants the data in S3 within 60 days. What is the right approach?
Q4. A regulated financial institution needs to migrate 200 TB of data to AWS, but policy prohibits shipping storage media off-site. What should the architect recommend?
Q5. A company has 300 TB of historical data plus 50 GB of new data generated daily. The network link is capped at 500 Mbps. What is the recommended migration pattern?
Q6. A mining company needs to run machine learning inference on video feeds at a remote site with intermittent connectivity, then ship the results to AWS. Which device is designed for this?
Q7. A team is migrating a large dataset and wants to know how the data on a Snowball is protected if the device is lost in transit. What is the correct description of the security model?
Q8. A company needs to move 100 PB of data to AWS in a single engagement. Which Snow Family option is designed for this scale?
Q9. A team copies a running database's files onto a Snowball and ships it. After ingest, the target database fails integrity checks. What is the most likely cause?
Q10. A company needs continuous, scheduled synchronization of an on-prem NFS share to Amazon EFS. Which service is the right fit?
Q11. A data center operator needs a graphical tool to manage file transfers to a Snowball Edge device without writing scripts. What should they use?
Q12. A company has 5 TB of data to move to S3 and a healthy 1 Gbps internet link. What is the most cost-effective approach?
Q13. A company wants to migrate 2 PB of data and is cost-sensitive. Which approach is likely most cost-effective?
Q14. A team is planning a Snowball migration and wants to know what to monitor during the copy phase. What is the most useful signal?
Q15. A scenario describes a facility that needs persistent, low-latency access to on-prem data from AWS applications, with new data written continuously. Which service is the right fit?
Peek into Tomorrow
Everything in this lesson assumed the data eventually lands in an AWS Region, and that the only question was how to get it there. But a growing class of workloads cannot accept that assumption at all. If a hospital must keep patient data physically on-premises for regulatory reasons, or a factory floor needs single-digit-millisecond latency to industrial control systems, then shipping data to a Region and back is not a migration problem — it is an architectural disqualifier. The open question is whether AWS infrastructure can be extended to the workload's location rather than the other way around.
Tomorrow's topic answers that question with three different mechanisms that all push AWS closer to the workload, but for different reasons. Outposts extends AWS APIs into your own data center, which is the answer when a data residency requirement makes a Region unusable. Local Zones place AWS compute and storage in specific metro areas, which is the answer when latency matters but you do not want to own hardware. Wavelength embeds AWS compute at the telco 5G edge, which is the answer for ultra-low-latency mobile applications. The interesting part is that these three options look similar on a slide but are chosen for entirely different constraints, and the exam tests whether you can tell them apart.