On-premise to AWS migration
Moving workloads to AWS requires more than replicating servers. This guide explains how to assess dependencies, choose migration strategies, build a secure foundation, and execute a tested cutover.
What an on-premise to AWS migration actually involves
An on-premise to aws migration moves applications, data, and operational responsibilities from organization-managed infrastructure into Amazon Web Services. The technical transfer is only one part: successful programs also redesign access controls, network connectivity, recovery procedures, financial ownership, and support responsibilities.
For decision-makers, the central question is whether migration produces a defensible improvement in resilience, delivery speed, capacity flexibility, or total cost. For practitioners, it is whether each workload can function correctly under different latency, identity, storage, and failure conditions.
This guide focuses on moving existing enterprise workloads into AWS, including virtual machines, databases, file services, and their dependencies. The objective is not to make everything cloud-native immediately. It is to choose the smallest safe change that meets a measurable business goal.
Define outcomes and constraints before selecting services
“Exit the data center” is a deadline, not a complete success criterion. Translate the business case into requirements that architecture and migration teams can verify.
Document at least:
- Exit scope: Which facilities, racks, contracts, and shared services must disappear?
- Service continuity: What downtime and data loss can each business process tolerate?
- Performance: What response times, transaction throughput, and batch completion windows are required?
- Compliance: Where may data reside, who may access it, and what evidence must be retained?
- Financial boundaries: What migration budget, dual-running period, and steady-state operating cost are acceptable?
- Operational ownership: Who will patch, monitor, recover, and pay for each workload?
Separate recovery time objective (RTO) from recovery point objective (RPO). RTO concerns service restoration time; RPO concerns tolerable data loss. Neither automatically determines the permitted migration outage, which needs its own agreement.
Define acceptance gates before implementation. For example, require representative load tests, a demonstrated backup restore, security approval, and an identified on-call owner before production cutover.
Step 1: Discover workloads and map dependencies
Start with an inventory that connects infrastructure to business services. A spreadsheet of virtual machines is insufficient if nobody knows which application depends on a particular database or identity server.
Collect:
- Server operating systems, versions, CPU, memory, and storage utilization.
- Database engines, extensions, sizes, growth, and replication capabilities.
- Network connections, DNS dependencies, firewall rules, and outbound integrations.
- Authentication flows, certificates, service accounts, and embedded credentials.
- Software licenses, hardware dependencies, support status, and application owners.
- Backup methods, recovery procedures, scheduled jobs, and maintenance windows.
Use existing VMware vCenter inventory, configuration management databases, monitoring platforms, and network-flow records. Tools such as Device42, ServiceNow Discovery, and AWS Migration Evaluator can contribute inventory, dependency, or cost information, depending on scope and supported collection methods.
Observe representative business cycles, not just an idle afternoon. Month-end reporting, overnight ETL, and seasonal peaks can expose requirements that average utilization conceals.
Turn dependencies into migration units
Group components that must move together into a migration unit. An application server and database communicating through frequent synchronous calls may perform poorly when split across a WAN.
Classify dependencies as:
- Move together: Latency-sensitive or tightly coupled components.
- Temporarily bridge: Dependencies that can tolerate hybrid connectivity.
- Replace first: Shared services that need an AWS-ready alternative.
- Retire: Components confirmed unused by both telemetry and owners.
Validate discovered traffic with application teams. A connection map shows communication, but rarely explains whether it is business-critical.
Step 2: Choose a strategy for each workload
AWS describes migration choices through its seven migration strategies. Use them as workload-level decisions rather than declaring one strategy for the entire estate.
| Strategy | Appropriate criteria | Main trade-off |
|---|---|---|
| Rehost | Tight exit deadline; application changes are risky | Moves existing inefficiencies and operational work |
| Replatform | Managed services offer value with limited application changes | Requires compatibility and operational testing |
| Refactor/re-architect | Scaling, resilience, or delivery constraints justify redesign | Highest engineering effort and scope risk |
| Repurchase | A SaaS product can replace existing functionality | Data conversion, integration changes, and vendor dependence |
| Relocate | A supported platform can move with limited architectural change | Platform economics, availability, and licensing need validation |
| Retain | Latency, regulation, hardware, or economics favor staying | Preserves hybrid dependencies and operating costs |
| Retire | The workload has no continuing business value | Requires retention, ownership, and shutdown checks |
A practical program might rehost a stable application onto Amazon EC2, replatform its database onto Amazon RDS later, and retire its obsolete reporting server immediately.
Avoid combining infrastructure migration, database-engine conversion, application decomposition, and identity replacement into one cutover unless the business case demands it. Each additional change complicates fault isolation and rollback.
Step 3: Model AWS costs and the transition budget
Compare total operating cost, not an EC2 estimate against a hardware purchase price.
Build a workload-level estimate using the AWS Pricing Calculator. Include:
- EC2 instances, EBS capacity, provisioned performance, and snapshots.
- Databases, replicas, backups, and high-availability configurations.
- Load balancers, NAT gateways, public IPv4 addresses, and data transfer.
- Logs, metrics, security services, and retention.
- VPN or Direct Connect connectivity and carrier charges.
- Software licensing, support, and specialist migration labor.
Model normal demand, peak demand, and disaster-recovery requirements separately. Rightsize using measured utilization, but preserve headroom for concurrency, memory pressure, and storage latency.
Account explicitly for dual running. Source infrastructure, AWS resources, replication, connectivity, and staffing may all remain chargeable during the transition. Data-center costs only disappear when contracts, facilities, and shared dependencies can actually be retired.
Delay long-term commitments until the target footprint stabilizes. Savings Plans and Reserved Instances can improve eligible workload economics, but premature purchases can lock spending to assumptions that change after migration.
Step 4: Build the AWS landing zone
A landing zone is the governed environment into which workloads migrate. Build it before production replication, not after servers arrive.
Establish account and identity boundaries
Use AWS Organizations to separate production, nonproduction, security, and log-archive responsibilities. AWS Control Tower can establish and govern a multi-account environment; its landing zone documentation explains the underlying approach.
Connect workforce access through IAM Identity Center and an appropriate identity provider. Require strong authentication, use temporary credentials, and avoid shared administrator accounts.
Service control policies set permission boundaries; they do not grant access. Test them against deployment, backup, and emergency-access procedures.
Design networking and operational controls
Allocate nonoverlapping address ranges for AWS and remaining on-premises networks. Resolve DNS forwarding, route propagation, inspection paths, and egress before moving workloads.
AWS Site-to-Site VPN is often useful for initial connectivity. Direct Connect can provide more predictable private connectivity, but requires provisioning lead time and a separate resilience design. Direct Connect is not inherently end-to-end encrypted; choose encryption controls according to the connection model and security requirements.
Implement centralized CloudTrail collection, appropriate AWS Config coverage, GuardDuty, encryption policies, backup policies, and cost-allocation tags. Capture infrastructure in Terraform, AWS CloudFormation, or AWS CDK so environments remain reproducible.
Step 5: Select migration tools by workload and data pattern
Different tools solve different transfer problems. Select based on compatibility, change rate, cutover requirements, and recovery behavior.
Servers and applications
AWS Application Migration Service, commonly called AWS MGN, supports replication and conversion for supported source servers into EC2. It is useful for rehosting, but does not automatically modernize applications or validate business behavior.
Check supported operating systems, replication prerequisites, target instance compatibility, and licensing. Test boot configuration, attached volumes, scheduled services, endpoint security agents, and monitoring after conversion.
Databases
AWS Database Migration Service (AWS DMS) can perform full loads and ongoing change replication for supported database combinations. Continuous replication can shorten the final outage, but does not eliminate cutover coordination.
For heterogeneous migrations, schema conversion and application changes are separate concerns. AWS DMS Schema Conversion or AWS Schema Conversion Tool may help where supported. Review stored procedures, data types, collations, extensions, and transaction semantics.
Native database replication or backup-and-restore may be simpler for a same-engine migration. Choose based on tested behavior rather than assuming a general-purpose migration service is always preferable.
Files and object data
AWS DataSync can move supported file and object datasets into services such as Amazon S3, Amazon EFS, and Amazon FSx.
Choose the destination by access semantics. S3 is object storage, not a drop-in replacement for an application expecting a POSIX filesystem or SMB share. Check permissions, metadata, locking behavior, throughput, and small-file performance.
Estimate transfer duration from usable throughput and dataset size, then allow for overhead and ongoing changes. Replication cannot converge if the source changes faster than the transfer path can sustain.
Step 6: Run a pilot and organize migration waves
Choose a pilot that is low-risk but representative. A trivial standalone server will not validate a program dominated by complex database-backed applications.
The pilot should exercise:
- Identity integration and hybrid DNS.
- Replication, deployment automation, and security controls.
- Application performance under representative load.
- Monitoring, backup restoration, and incident response.
- Cutover communication and rollback decisions.
- Actual AWS spending against the estimate.
Use the results to update reusable runbooks and infrastructure templates.
Build subsequent waves around dependency groups, business calendars, and support capacity. Avoid moving every application owned by one operations team simultaneously. A wave is only manageable if engineers can investigate failures without abandoning other critical services.
Give each workload explicit gates: discovery complete, target ready, test passed, cutover approved, stabilization complete, and source retirement authorized.
Step 7: Rehearse cutover and define rollback limits
A cutover runbook should identify the action, owner, expected duration, verification, and failure response for every step.
A typical sequence is:
- Confirm approvals, support coverage, and the change freeze.
- Verify target capacity, backup status, replication health, and observability.
- Pause source writes or place the application into maintenance mode.
- Allow final replication to complete and verify consistency.
- Update endpoints, routing, or DNS as planned.
- Enable the target application and run business-critical checks.
- Monitor errors, latency, queues, integrations, and data integrity.
- Declare success only after the agreed observation period.
Lowering DNS TTL in advance can help, but clients and connection pools may retain old endpoints. Include reconnection behavior in rehearsals.
Rollback is not simply switching DNS back. Once AWS accepts new writes, the source may be stale. The plan must explain whether reverse replication, reconciliation, or recovery from backup is possible—and when rollback stops being safe.
Prefer measurable triggers, such as failed payment processing or a sustained breach of an agreed latency threshold, over “roll back if performance seems bad.”
Step 8: Stabilize, optimize, and decommission
During stabilization, compare AWS behavior with the source baseline. Review transaction success, tail latency, batch duration, replication lag, resource saturation, and support incidents.
Then optimize deliberately:
- Rightsize compute and database resources.
- Remove temporary replication and test infrastructure.
- Tune storage performance and retention policies.
- Review cross-AZ and outbound data-transfer patterns.
- Replace temporary administrator permissions with scoped roles.
- Confirm backups through restoration, not successful-job notifications alone.
Decommission source systems only after workload-owner approval and retention checks. Preserve required records, remove credentials, update inventories, terminate licenses where appropriate, and securely dispose of storage.
Track realized benefits against the original case. If data-center spending remains unchanged, investigate shared dependencies and contractual exit conditions rather than declaring savings from an infrastructure move alone.
Common mistakes that undermine AWS migrations
- Copying server sizes unchanged: Historical allocations often exceed real demand. Measure first, then validate smaller targets under load.
- Leaving chatty dependencies on-premises: WAN latency can overwhelm otherwise adequate compute. Move tightly coupled components together.
- Treating replication as backup: Replication may reproduce deletion or corruption. Maintain independent, recoverable backups.
- Assuming managed means maintenance-free: RDS still requires capacity planning, access control, upgrade decisions, and recovery testing.
- Ignoring license mobility: Validate Microsoft, Oracle, and other vendor terms before committing to the target design.
- Skipping ownership handover: A running application without support coverage, alerts, and recovery procedures is not migration-complete.
Frequently asked questions
How long does an on-premise to AWS migration take?
A small, well-understood workload may move in weeks, while a large estate can require many months or longer. Dependency discovery, procurement, database compatibility, compliance reviews, and data-center exit conditions often determine the schedule more than transfer speed.
Can we migrate to AWS without downtime?
Some architectures support near-zero-downtime transitions using replication and controlled traffic switching. Many existing applications still need a short write freeze. The answer depends on transaction consistency, replication capabilities, client behavior, and whether the application safely supports simultaneous environments.
Is rehosting cheaper than replatforming?
Rehosting generally demands less initial application engineering, but may preserve expensive licensing and manual operations. Replatforming can reduce ongoing work while adding migration and compatibility costs. Compare both over a defined planning horizon, including labor, resilience, and exit costs.
What should move to AWS first?
Start with a workload that has clear ownership, understood dependencies, recoverable data, and limited business impact—but still exercises the target architecture. Avoid both the most critical application and an unrepresentative toy system. Use the pilot to prove the migration process before increasing wave complexity.
For related platform-transition and modernization planning, browse more Migration topics.
Ask the community and get answers from practitioners.