AWS Application Discovery Service — Assessment & Portfolio Analysis
Recap: From Classification to Evidence
Day 43 gave us the vocabulary for the whole phase: the 7 Rs, with Rehost mapped to MGN, Relocate mapped to VMware Cloud on AWS, and Replatform described as lift-tinker-and-shift — the self-managed MySQL to RDS move being the canonical example. Migration Hub sat above all of it as the tracking surface, a single place to watch application status move from discovery through cutover. That framework was a prerequisite for what we do here, and the dependency runs in a specific direction: the 7 Rs tell you what decision to make, but they say nothing about how you justify it. Assigning an application to Rehost is a judgment call, and a judgment call made in a steering committee without data is how migration programs end up with target instance sizes picked from a spreadsheet of guesses.
Application Discovery Service is the evidence layer underneath the classification. It is what turns "we think this app is a Rehost candidate" into "this app runs on four hosts, talks to two databases and one internal API, and its peak CPU utilization is 12%." Migration Hub tracking progress is only meaningful if the plan it tracks was built from measured inventory rather than from a CMDB that was last reconciled three years ago.
Foundations You'll Need Today
This day assumes a handful of ideas that are second nature to someone who has run a data center but are never spelled out in the AWS documentation, because the documentation assumes you already have them. Here they are, in the order the day leans on them.
The CMDB, and why nobody trusts it
A configuration management database, or CMDB, is the spreadsheet-to-database evolution of an IT department's answer to a simple question: what do we actually own and run? Every server, every application, every database, every network device gets a record, and those records are supposed to be updated whenever something changes. In practice they are updated by hand, by people who are busy, and the updates lag reality by months or years. The result is a document that everyone references and nobody believes. This matters because migration planning is fundamentally an inventory problem: you cannot decide what to move, in what order, or onto what size of machine until you know what exists. The CMDB is the thing you would use if it were accurate, and Application Discovery Service exists because it usually is not.
Virtual machines, hypervisors, and vCenter
A virtual machine is a computer that is not a physical computer — it is a software-defined machine running inside a real one, sharing that real machine's CPU, memory, and disk with other virtual machines. The software that makes this possible is called a hypervisor, and it is what sits between the physical hardware and all the virtual machines running on it. VMware is the dominant commercial hypervisor in enterprise data centers, and vCenter is VMware's management layer: a central service that knows about every virtual machine on every host it manages, what each one is configured with, and how much CPU and memory each is using. The important consequence for today is that vCenter already holds a rich inventory of everything virtualized, and it exposes that inventory through an API — a programmatic interface that other software can query. That is the entire reason an agentless discovery appliance can work: it does not need to touch the virtual machines at all, because vCenter will happily tell it everything about them.
Agents, and why installing one is a project
An agent is a small piece of software installed on a machine that runs continuously and reports back to a central service. Monitoring tools use agents, backup tools use agents, security tools use agents. The advantage of an agent is depth: because it runs inside the operating system, it can see things that are invisible from outside, like which processes are running and what network connections those processes are making. The disadvantage is that installing it means changing a production server. In most enterprises that triggers change control — a formal process where a proposed change is reviewed, scheduled, and approved before anyone touches the machine. Multiply that by hundreds or thousands of servers and the agent rollout becomes a project with its own timeline, its own maintenance windows, and its own risk of slipping. This is the tradeoff at the heart of today's topic: agents give you better data, and they cost you operational effort to deploy.
Right-sizing, and why one CPU reading is worthless
Right-sizing means choosing a target machine size that matches what the workload actually needs, rather than copying the size of the machine it currently runs on. The reason this is not trivial is that on-premises servers are chronically over-provisioned — a server was bought with eight cores because that was the smallest sensible purchase at the time, and it has run at 10% utilization ever since. If you migrate it to an eight-core cloud instance, you have moved the waste rather than fixed it. To right-size properly you need utilization data, and specifically you need it over a long enough window to see the peaks. A server that looks idle on a Tuesday might run a month-end financial close that saturates it for six hours. A single reading, or a reading taken during a quiet week, will tell you the server is small when it is not, and the mistake only surfaces in production at the first peak. This is why discovery collects continuously over a window rather than sampling once.
Dependency maps and migration waves
A dependency is a connection between two systems where one needs the other to function — an application server that queries a database, a web tier that calls an internal API, a service that checks a license server on startup. A dependency map is a picture of all those connections across an estate. It matters because migrations are executed in waves: you move a group of applications together, verify they work, then move the next group. The grouping is not arbitrary. If application A depends on application B and you migrate A in wave one and B in wave three, then A is broken for the duration between the two waves, because the thing it needs is still sitting in the old data center. Wave sequencing is therefore the problem of finding clusters of applications that depend mostly on each other and cutting them over together — and you cannot find those clusters without a dependency map. This is the single most important reason discovery exists, and it is why the agent-versus-agentless decision later in this day is a real decision rather than a preference.
With that grounding, here is why Application Discovery Service exists and what problem it actually solves.
1. Why This Is on the Exam
Application Discovery Service appears on SAP-C02 in the migration and modernization domain, and it almost never appears as the answer to a question about discovery itself. It appears as the first step in a multi-part scenario where the constraint is that the customer does not actually know what they have. The exam likes this shape because it separates candidates who reach for a migration tool immediately from candidates who recognize that you cannot size a target, sequence a wave, or choose between Rehost and Replatform until you have measured the source estate.
The architectural problem it solves is the inventory gap. Most enterprises pursuing a data center exit have a configuration management database, and most of those CMDBs are partially wrong. They were populated by hand, updated inconsistently, and they describe what someone believed was running rather than what is actually running. They also almost never capture the thing that matters most for migration sequencing: which servers talk to which other servers, and how much. A CMDB entry for an application server does not tell you that it makes 400 calls a minute to a licensing daemon on a host nobody remembers installing. Discovery exists to close that gap with measured data.
The second reason it is exam-relevant is right-sizing. A migration that lifts a fleet of eight-core, 32 GB virtual machines into identically sized EC2 instances is technically successful and financially indefensible. Discovery collects utilization data — CPU, memory, disk — over a window, and that data is what feeds the target instance recommendation. The exam will describe a customer with a large on-prem footprint, a fixed migration deadline, and a mandate to reduce run-rate cost, and the correct first move is discovery, not MGN.
There is also a governance dimension. Discovery data is what you present to a steering committee to justify wave sequencing. It is what you use to identify applications that are Retire candidates because nothing has connected to them in ninety days. It is what tells you that two applications you assumed were independent share a database, which means they cannot be migrated in separate waves without a temporary cross-environment dependency. None of that is derivable from an org chart.
Finally, the exam tests the boundary between discovery and assessment. Discovery is inventory and dependency collection. Assessment is the analysis layer on top — Migration Evaluator produces a business case with projected costs, and Migration Hub Strategy Recommendations produces modernization path recommendations. Candidates who conflate these will pick the wrong tool when a question asks specifically for a cost projection or a recommended target platform.
2. How Discovery Actually Works
Application Discovery Service has two collection mechanisms, and the difference between them is the single most important mechanical fact about the service. The agentless option is the Discovery Connector, a virtual appliance you deploy into your VMware environment as an OVA. It talks to vCenter over the VMware API and collects virtual machine inventory — hostnames, IP addresses, MAC addresses, CPU and memory allocation, disk configuration, and basic performance counters that vCenter already tracks. It requires no software on the guest operating systems, which means it works on machines you cannot touch, machines running operating systems you do not support, and appliances whose vendors will not permit agent installation.
The agent-based option is the Discovery Agent, a lightweight collector installed on each server. It reports the same inventory data plus two things the connector cannot see: process-level detail and network connections. Specifically, the agent records which processes are running, their resource consumption, and the outbound network connections each process makes — the remote address, port, and connection frequency. That connection data is what produces the dependency map. The connector can tell you a VM exists and how much CPU it uses; only the agent can tell you that the process on that VM is talking to a specific database on a specific host.
Both mechanisms write into the same place: the discovery data store, which is per-region and per-account, and which is queried through the Discovery API or viewed in the Migration Hub console. The data model is roughly three layers. The first is the server record — one per discovered host, with its configuration and utilization. The second is the process record, populated only by agents, describing what runs on each server. The third is the connection record, also agent-only, describing observed network flows between processes on different servers. The dependency map is a rendering of that third layer.
Collection is continuous rather than one-shot. Agents report on an ongoing basis, so utilization data accumulates into a window you can analyze rather than a single sample. This matters because a single CPU reading is nearly useless for right-sizing — you need to see the peak, and you need to see whether the peak is a daily pattern or a monthly batch job. The console exposes utilization over time, and the export path lets you pull the raw data into your own analysis if you want to model it differently.
Data can be exported. The Discovery API supports exporting collected data to S3, either as a one-time export or on a scheduled basis, and the export includes the server, process, and connection records. This is the mechanism by which discovery data reaches Migration Hub Strategy Recommendations and Migration Evaluator, and it is also how you get the data into a spreadsheet or a BI tool when you need to argue about wave sequencing with people who do not log into the AWS console.
One mechanical detail worth internalizing: the agent communicates outbound to the AWS endpoint over TLS, typically through a proxy or NAT path, and it does not require inbound connectivity. That is why discovery can be deployed in a segmented network where the migration team has no ability to open inbound firewall rules. The connector similarly needs outbound access to the AWS endpoint and network access to vCenter.
3. The Core Decision Boundary: Agent or Agentless
Every discovery scenario reduces to one fork: do you need dependency data, or do you only need inventory and utilization? That question determines whether you deploy agents, and it is the fork the exam tests most often. The reason it is a genuine fork rather than a preference is that the two mechanisms produce materially different datasets, and the difference is not a matter of degree. Agentless discovery gives you a complete picture of what exists. Agent-based discovery gives you a complete picture of what exists plus how it is connected. If your migration plan requires wave sequencing — and any migration of more than a handful of applications does — you need the connection data, because wave sequencing is fundamentally the problem of grouping applications that talk to each other into the same migration window.
The counter-pressure is operational. Agents require installation on every server, which means change control, which means a maintenance window per server or a coordinated rollout, which in a large estate can take weeks. Agents also require an operating system the agent supports, and they require someone with administrative access to install them. In environments with strict change management, a large population of appliances, or operating systems outside the supported matrix, agentless is the only option that can start today.
The practical resolution most organizations land on is a hybrid: deploy the connector first to get immediate inventory coverage across the whole VMware estate, then deploy agents selectively on the servers that matter for dependency mapping. The connector gives you the denominator — how many servers exist — while agents give you the numerator for the applications you are actually planning to move first. This is also the pattern the exam tends to reward, because it acknowledges the real constraint that you cannot always instrument everything.
| Capability | Discovery Connector (agentless) | Discovery Agent |
|---|---|---|
| Deployment | OVA appliance in VMware, talks to vCenter | Per-server software install |
| VM inventory (hostname, IP, CPU, memory, disk) | Yes | Yes |
| Utilization over time | Yes, from vCenter counters | Yes, measured in-guest |
| Running processes | No | Yes |
| Network connections between servers | No | Yes |
| Dependency map | Not produced | Produced from connection data |
| Works on unsupported OS / appliances | Yes | No |
| Change control burden | One appliance | Every server |
| Non-VMware sources | No (VMware only) | Yes, any supported OS |
Note the last row carefully. The connector is VMware-specific because it depends on the vCenter API. If the source estate includes physical servers, Hyper-V, or another hypervisor, the connector cannot see them at all, and agents are the only path. Exam scenarios that mention physical servers alongside virtual ones are usually testing whether you noticed this.
4. Collection Modes and Their Tradeoffs
Beyond the agent-versus-agentless choice, there are configuration decisions that change what the data looks like and how much work the analysis downstream becomes. The first is the collection window. Utilization data is only as good as the period it covers, and the period has to be long enough to capture the workload's real peak. A server that runs a month-end financial close will look idle if you collect for a week in the middle of the month, and right-sizing it from that data will produce an instance that falls over on the first of the month. The general guidance is to collect across at least one full business cycle, which for most organizations means a month, and longer if there are quarterly or annual peaks.
The second decision is scope. Discovery can be deployed broadly across an entire data center or narrowly against a specific application's server group. Broad deployment gives you the complete dependency graph, which is what you need to detect the cross-application couplings that break wave plans. Narrow deployment is faster to stand up and produces less data to analyze, but it will miss dependencies that leave the scoped group — and those are exactly the dependencies that cause cutover surprises. The failure mode of narrow scoping is discovering during cutover that the application you moved calls a service you did not move.
The third is the export and retention pattern. Discovery data lives in the service, and you can export it to S3 on a schedule. Organizations that intend to do serious analysis — building a dependency graph in a graph database, joining discovery data against CMDB records to find discrepancies, or feeding a cost model — generally set up scheduled exports early rather than trying to pull everything at the end. The export is also the integration point with Migration Hub Strategy Recommendations, which consumes discovery data to produce modernization path recommendations.
The fourth is how you handle the agent rollout itself. Installing agents at scale is an operational project, not a configuration toggle. The realistic approaches are to bundle the agent into an existing configuration management pipeline, to deploy it via a scheduled task pushed by Group Policy or an equivalent, or to accept a slower manual rollout prioritized by migration wave. The choice affects your timeline more than your data quality, but it is the step that most often slips in real programs.
There is a fifth consideration that is easy to overlook: the agent's own footprint. It is lightweight by design, but it is still a process on a production server, and in environments with strict performance or compliance requirements that fact needs to be socialized before rollout rather than discovered during it. The same applies to the connector appliance, which consumes vCenter API calls and needs to be sized appropriately for the number of VMs it is polling.
5. Sizing, Limits and Quotas
The numbers that matter for discovery are mostly about scale and about what the service will and will not do at volume. The connector is deployed as a virtual appliance and is sized against the number of VMs in the vCenter it is polling; a single connector has a practical ceiling, and larger estates require multiple connectors, typically one per vCenter or per large cluster. The agent is designed to be low-impact and runs as a background process, reporting on an interval rather than continuously streaming.
Discovery data is regional. The service collects into the region where you have configured it, and the data store is per-account. This matters for organizations with data residency constraints, because it means you can choose the collection region deliberately rather than having it chosen for you. It also means that if you are running a multi-region migration, you need to think about where discovery data lives relative to where the migration tooling runs.
Export to S3 is the escape hatch for anything the console does not render well. The export produces structured records for servers, processes, and connections, and once in S3 you can query it with Athena, load it into a warehouse, or process it with whatever tooling your organization already has. There is no meaningful limit on how much you can export; the limit is on how much you can usefully analyze.
| Dimension | What to know |
|---|---|
| Connector deployment unit | One OVA appliance per vCenter (or per large cluster); multiple connectors for large estates |
| Agent deployment unit | One agent per server, installed in-guest |
| Data store scope | Per-region, per-account |
| Collection cadence | Continuous; utilization accumulates over a window rather than a single sample |
| Recommended collection window | At least one full business cycle (typically a month) to capture real peaks |
| Export target | S3, one-time or scheduled, covering server/process/connection records |
| Dependency data availability | Agent-based collection only |
| Non-VMware source support | Agent-based collection only |
One quota-adjacent point that shows up in exam scenarios: discovery is not a real-time monitoring system and should not be treated as one. It is not a replacement for CloudWatch, and it does not produce alarms. Its output is a dataset for planning, refreshed on a cadence appropriate to planning rather than to incident response. Candidates who describe discovery as a monitoring solution have misread the service.
6. Failure Modes and What They Look Like in Production
The most common failure mode is incomplete dependency data, and it is insidious because it does not announce itself. You deploy agents on the servers you know about, the dependency map renders, and it looks plausible. What it does not show is the dependency on a server you did not instrument — a licensing server, a jump host, a shared file server, an appliance. The map is not wrong; it is incomplete, and incompleteness in a dependency map is indistinguishable from absence of dependency unless you know the denominator. This is the strongest argument for deploying the connector across the whole estate first: it establishes how many servers exist, so you can tell whether your agent coverage is 95% or 40%.
The second failure mode is stale or misleading utilization data. If the collection window is too short, or if it happens to fall during a quiet period, the right-sizing recommendation will be too small. The symptom appears months later as an instance that performs fine in staging and falls over in production at the first peak. The diagnostic move is to check the collection window against the workload's known cycle — if the application has a monthly or quarterly peak, a two-week window is not evidence.
The third is agent installation failure on a subset of servers, which is usually silent. The agent installs, the service starts, and then the agent cannot reach the AWS endpoint because of a proxy configuration or a firewall rule that was not opened for that subnet. The server appears in inventory — because the connector saw it — but has no process or connection data. The symptom is a server that shows up in the server list but contributes nothing to the dependency map. The first diagnostic move is to check agent connectivity from the host, not to reinstall the agent.
The fourth is connector saturation. A single connector polling a very large vCenter can fall behind, producing inventory that lags reality. The symptom is servers that were decommissioned weeks ago still appearing, or newly provisioned servers not appearing. The fix is additional connectors, and the diagnostic is comparing connector-reported inventory against a known-good source.
The fifth, and the one that causes the most program-level damage, is treating discovery output as a migration plan. Discovery tells you what exists and how it is connected. It does not tell you what order to migrate in, which applications should be consolidated, or which should be retired. Those are analysis decisions made on top of the data. Programs that skip the analysis step and treat the dependency graph as a wave plan end up with waves that are technically valid and operationally nonsensical.
7. The Operational and SRE Angle
Discovery is a planning tool, not a production system, so the SRE angle is about the discovery pipeline itself rather than about the workloads it observes. The pipeline has three components that can fail independently: the connector appliance, the agents, and the export path to S3. Each needs its own health check, and none of them is covered by the alarms you already have on production workloads.
For the connector, the health signal is whether it is successfully polling vCenter and whether the inventory it reports is current. A connector that has lost its vCenter credentials will keep running and keep reporting nothing new, which looks identical to a stable environment. The runbook shape is a periodic reconciliation: compare the connector's server count against vCenter's VM count, and alert on divergence beyond a threshold. That single check catches credential expiry, network partition, and saturation.
For agents, the health signal is the last-report timestamp per server. An agent that has stopped reporting is either stopped, disconnected, or the host is down. In a migration program, a server whose agent stopped reporting two weeks ago is a server whose dependency data is two weeks stale, and if that server is in the next wave, the wave plan is built on stale evidence. The runbook is a weekly report of servers with agent data older than a threshold, routed to the migration team rather than to on-call, because it is a planning defect rather than an incident.
For the export path, the health signal is the presence and freshness of the S3 objects. A scheduled export that silently stops produces a downstream analysis that quietly goes stale. This is the classic silent-failure pattern: nothing alarms, the dashboard still renders, and the numbers are three weeks old. The check is a freshness assertion on the export prefix, and it belongs in whatever monitoring you use for data pipelines.
There is a broader SRE point here that is worth stating explicitly. Discovery data has a shelf life. A dependency map built six months ago describes an environment that has changed, and the rate of change in most enterprises is high enough that a six-month-old map is a liability rather than an asset. If discovery is a one-time project that ends when the migration plan is written, the plan will be executed against a stale picture. The operational posture that works is continuous collection for the duration of the migration program, with the dependency map treated as a living artifact that is re-examined before each wave.
8. Edge Cases and Exam Gotchas
The single most-tested gotcha is the agentless-versus-agent dependency distinction. If a question asks for a dependency map, process-level detail, or network connections between servers, the answer requires agents. If it asks for inventory, utilization, or right-sizing data and mentions constraints that make agent installation impractical, the answer is the connector. Questions that mention both needs are usually testing whether you propose the hybrid approach.
The second gotcha is the VMware-only limitation of the connector. Any scenario that includes physical servers, a non-VMware hypervisor, or servers in a colocation facility without vCenter cannot be covered by the connector alone. This is a common distractor setup: the scenario describes a mixed estate and offers the connector as an answer, and the correct answer is agents or a hybrid.
The third is the distinction between discovery and assessment. Discovery collects data. Migration Evaluator builds a business case with cost projections. Migration Hub Strategy Recommendations produces modernization path recommendations. A question that asks for a projected cost of migrating a portfolio is asking for Migration Evaluator, not discovery, even though discovery data feeds it.
The fourth is the collection window. Any scenario that describes a workload with a periodic peak and a short collection window is testing whether you noticed that the utilization data is not representative. The correct answer usually involves extending collection rather than accepting the right-sizing recommendation.
The fifth is the relationship to Migration Hub. Discovery feeds Migration Hub; it does not replace it. Migration Hub is the tracking surface where applications move through states. Discovery is one of the inputs that populates it. A question that asks where you track migration progress across multiple tools is asking about Migration Hub.
The sixth is the data residency angle. Discovery data is regional and per-account. In a scenario with regulatory constraints on where data can be stored, the collection region is a deliberate choice, and the exam may test whether you considered it.
The seventh, and the one most often missed, is that discovery does not produce a migration plan. It produces the inputs to one. Any answer option that describes discovery as generating wave plans, recommending target platforms, or sequencing migrations has overstated the service.
9. Discovery vs. the Services It Gets Confused With
The services adjacent to discovery fall into three groups: the assessment layer above it, the migration execution layer below it, and the general-purpose inventory tools that look similar but serve a different purpose. Getting the boundaries right is most of what the exam tests in this area, because the individual services are not complicated — the confusion is entirely about which layer a given question is asking about.
The assessment layer is Migration Evaluator and Migration Hub Strategy Recommendations. Migration Evaluator is the business case tool: it takes discovery data and produces a projected cost comparison between running on-premises and running on AWS, including licensing considerations. Strategy Recommendations is the modernization advisor: it takes discovery data and produces recommendations about which applications are candidates for rehost, replatform, or refactor, along with target platform suggestions. Both consume discovery output; neither collects it.
The execution layer is MGN for rehost and DMS for database migration. These are the tools that actually move things. They come after discovery and after the plan, and they do not collect inventory themselves — MGN replicates servers it has been told about, and DMS replicates databases it has been configured for. A question that asks how to move 200 servers to EC2 is asking about MGN; a question that asks how to determine which 200 servers to move is asking about discovery.
The general-purpose inventory tools are AWS Config and Systems Manager Inventory. Config records resource configuration and configuration changes for resources already in AWS, and it is a compliance and drift tool. Systems Manager Inventory collects metadata from managed instances, which by definition are already in AWS and already running the SSM agent. Neither can see an on-premises estate, which is the entire point of Application Discovery Service. A scenario about discovering on-premises servers cannot be answered with Config or SSM Inventory.
| Service | Layer | What it does | Pick it when… |
|---|---|---|---|
| Application Discovery Service | Discovery | Collects on-prem inventory, utilization, processes, and network connections | You need to know what exists and how it is connected before planning |
| Migration Evaluator | Assessment | Builds a cost business case from discovery data | You need a projected on-prem vs. AWS cost comparison |
| Migration Hub Strategy Recommendations | Assessment | Recommends modernization paths and target platforms | You need guidance on rehost vs. replatform vs. refactor per app |
| Migration Hub | Tracking | Tracks application migration status across tools | You need one place to watch progress across the program |
| AWS MGN | Execution | Replicates and cuts over servers to EC2 | You have decided to rehost and need to move the servers |
| AWS DMS | Execution | Migrates database data, with CDC for minimal downtime | You are moving database contents, not servers |
| AWS Config | Governance | Records configuration and changes for AWS resources | You need compliance and drift detection inside AWS |
| SSM Inventory | Operations | Collects metadata from managed EC2/on-prem instances with SSM agent | You need operational metadata for instances already under management |
The rule of thumb that resolves most of these: if the question is about an estate you cannot see, the answer is Application Discovery Service. If it is about deciding what to do with an estate you can now see, the answer is in the assessment layer. If it is about moving something you have already decided to move, the answer is in the execution layer. And if it is about resources already in AWS, the answer is almost never discovery.
Hands-on Lab: Building a Discovery Dataset You Can Defend (45 min)
This lab walks through the full discovery workflow against a small simulated estate, with the emphasis on the difference between what agentless and agent-based collection actually produce. You will not need a real vCenter; the goal is to build the reasoning and the artifacts, and to see where the two collection modes diverge. Work through the steps in order and keep notes, because the analysis at the end is the part that matters.
Step 1 — Define the estate. Write down a fictional but realistic estate of twelve servers: four application servers, two database servers, one file server, one licensing server, one jump host, one monitoring server, one build server, and one appliance you cannot install software on. Assign each a role and a rough CPU/memory profile. The point of including the appliance is that it forces you to confront the coverage gap early.
Step 2 — Model agentless collection. For each of the twelve servers, write down what the Discovery Connector would report: hostname, IP, CPU allocation, memory allocation, disk configuration, and a utilization figure. Note that all twelve appear, including the appliance, because the connector sees VMs regardless of what runs inside them. Note also that you have no idea which servers talk to which.
Step 3 — Model agent-based collection. Now assume agents are installed on ten of the twelve servers — everything except the appliance and the jump host, which you could not get change approval for. For each of the ten, write down the processes running and the outbound connections. Deliberately include a connection from an application server to the licensing server, and a connection from two different application servers to the same database. Those two connections are the ones that will change your wave plan.
Step 4 — Build the dependency map. Draw the graph from the connection data in step 3. Then answer three questions in writing. First, which applications are coupled through shared infrastructure? Second, which servers have no observed inbound connections, and are therefore Retire candidates? Third, which servers appear in inventory but contribute nothing to the map, and what does that tell you about your coverage?
Step 5 — Identify the coverage gap. The appliance and the jump host have no agent data. Write down the specific risk this creates: if an application server depends on the appliance, you will not see it, and the dependency will surface during cutover. Then write down the mitigation — in this case, either getting agent approval for those two hosts or explicitly documenting them as known unknowns to be validated manually before the wave that contains their dependents.
Step 6 — Right-size from utilization. Take the utilization figures from step 2 and produce a target instance recommendation for each server. Then deliberately shorten the collection window to three days and redo the recommendation for the server you designated as having a month-end peak. Note how much smaller the recommendation becomes and what that would cost you in production.
Step 7 — Export and analyze. Describe the export you would configure: what records it contains, where it lands in S3, and what you would query it for. A reasonable answer includes joining the server records against the connection records to produce a per-application dependency count, which is the input to wave sequencing.
Step 8 — Write the one-page summary. Produce a single page that a steering committee could read: how many servers exist, how many are instrumented, how many applications are identified, how many cross-application dependencies were found, and what the coverage gap is. This artifact is the actual deliverable of discovery, and it is what makes the 7 Rs classification defensible rather than aspirational.
Scenario Question Drills (20 min)
Q1. A migration team needs to map network dependencies between individual processes on each server before planning application groupings for migration waves. What captures this?
Q2. An enterprise has 400 VMware VMs plus 60 physical servers in a colocation facility. They need a complete inventory before planning migration waves. What is the correct approach?
Q3. A customer wants a projected cost comparison between running their current data center and running the same workloads on AWS, including licensing. Which service produces this?
Q4. Discovery has been running for two weeks and the right-sizing recommendations look aggressive — several eight-core servers are recommended as two-core instances. The application has a month-end batch peak. What is the most likely problem?
Q5. A dependency map shows an application server with no observed inbound connections and near-zero utilization for ninety days. What does this most likely indicate?
Q6. A server appears in the discovery inventory but contributes no process or connection data to the dependency map. What is the first diagnostic step?
Q7. Which statement about the Discovery Connector is accurate?
Q8. A migration program wants to feed discovery data into a graph database to analyze application coupling. What is the mechanism?
Q9. A customer asks which service will recommend whether each application should be rehosted, replatformed, or refactored, with target platform suggestions. What should you recommend?
Q10. Discovery agents have been deployed on 40% of a 500-server estate. The dependency map looks complete and plausible. What is the primary risk?
Q11. Which of the following can Application Discovery Service NOT do?
Q12. A company has strict change control and cannot install software on 300 production servers within the migration timeline, but still needs inventory and utilization data to size targets. What should they deploy?
Q13. Discovery data was collected six months ago and the migration plan was written from it. The first wave is about to execute. What is the most important action?
Q14. A single Discovery Connector is polling a vCenter with 3,000 VMs. Servers decommissioned weeks ago still appear in inventory, and newly provisioned servers do not. What is the most likely cause?
Q15. Which pairing correctly matches a discovery input to the downstream service that consumes it?
Peek into Tomorrow
Discovery has given us a defensible inventory and a dependency map, and the 7 Rs have given us a classification for each application. What none of that addresses is the actual mechanics of the move. When an application is classified as Rehost, the question that follows immediately is how the server gets from a data center to EC2 without a multi-day outage, and how you validate that the target works before you commit to the cutover. That is a harder problem than it sounds, because the source server keeps running and keeps changing while you prepare the target.
The open question is whether you can replicate a running server continuously, in the background, without disrupting it, and then cut over in a window measured in minutes rather than days. That requires continuous block-level replication to a staging area in AWS, the ability to launch test instances from that staging area at any point without affecting the source, and a cutover step that is short enough to fit inside a maintenance window. It also raises a question about the tooling itself: the older CloudEndure and Server Migration Service workflows are being replaced, and knowing which tool is current matters for exam answers. Tomorrow's topic is the service that answers all of this.
Sources
- AWS Application Discovery Service — User Guide: What Is Application Discovery Service?
- AWS Application Discovery Service — Discovery Agent
- AWS Application Discovery Service — Discovery Connector
- AWS Application Discovery Service — Exporting Discovered Data
- AWS Application Discovery Service — How It Works
- AWS Migration Hub — User Guide
- AWS Migration Evaluator — Product Page
- Migration Hub Strategy Recommendations — User Guide
- AWS Prescriptive Guidance — Migration Discovery and Assessment