Monolith to microservices migration
Microservices solve specific scaling and delivery problems, but they also introduce distributed-systems costs. This guide explains how to choose service boundaries, migrate data safely, and measure whether each extraction is worth keeping.
Start with the problem, not the architecture
A monolith to microservices migration is a change to software architecture, data ownership, delivery workflows, and operational responsibility—not simply a project to split a codebase. For engineering leaders, the central question is whether independent services will remove measurable constraints. For practitioners, the challenge is preserving correctness while the old and new systems coexist.
Microservices can enable independent deployment, targeted scaling, and clearer ownership. They also introduce network failures, eventual consistency, more infrastructure, and harder debugging. A well-structured monolith may remain the better choice when one team can comfortably own the system.
The safest strategy is incremental: identify a costly constraint, extract a bounded capability, validate the benefits, and continue only when the evidence supports it.
Decide whether migration is justified
Before choosing Kubernetes or a messaging platform, document what the current architecture prevents you from doing.
Use concrete decision criteria
Look for recurring problems with identifiable architectural causes:
- Release coupling: Unrelated teams must coordinate deployments because their changes share one release artifact.
- Uneven resource demand: One capability consumes disproportionate CPU, memory, or database capacity, forcing the whole application to scale.
- Ownership conflicts: Several teams repeatedly modify the same modules without clear business ownership.
- Isolation requirements: A capability needs a distinct security boundary, availability target, or compliance treatment.
- Change bottlenecks: A business area changes frequently but must pass through broad regression testing and deployment coordination.
These symptoms require investigation. Slow queries, inadequate indexing, and unreliable tests rarely disappear when code becomes a network service.
Establish a baseline using deployment lead time, failed deployments, recovery time, tail latency, infrastructure spend, and operational workload. Define the improvement each proposed extraction should deliver.
Compare migration with simpler alternatives
| Situation | Preferred starting point | Main trade-off |
|---|---|---|
| Tangled modules within one team | Modular monolith with enforced boundaries | Less operational complexity, but deployments remain shared |
| One expensive background workload | Extract a worker or job processor | Targeted isolation without redesigning everything |
| Multiple autonomous domain teams | Incremental service extraction | Independent delivery requires mature operations |
| Unclear business boundaries | Domain discovery and internal refactoring | Slower visible progress, lower boundary mistakes |
| Coordinated releases caused by shared contracts | Contract and dependency redesign | More services alone will not remove coupling |
Do not use service count as a success metric. A migration that ends with three services and a smaller monolith may be more effective than one that produces dozens.
Define service boundaries around business capabilities
A service should own a cohesive business capability and its rules. “Orders,” “inventory reservations,” and “billing” can be useful boundaries; “controllers,” “database access,” and “validation” usually are not.
Domain-driven design provides a useful vocabulary: a bounded context defines where a particular model and its terminology apply. An order in fulfillment may have different responsibilities from an order in financial reporting.
Map dependencies before extracting code
Combine design workshops with runtime and repository evidence:
- Trace business workflows across modules, tables, scheduled jobs, and external integrations.
- Identify tables written by multiple modules.
- Review change history for modules that frequently change together.
- Map synchronous calls and transaction boundaries.
- Assign a team accountable for each proposed capability.
Tools such as ArchUnit for Java and NetArchTest for .NET can enforce dependency rules inside the monolith before extraction. OpenTelemetry instrumentation can expose runtime dependencies that static analysis misses.
Prefer boundaries that contain most business decisions locally. If every operation requires several synchronous calls back into the monolith, the proposed service may be too small or poorly bounded.
Build the operational foundation first
Independent deployment is valuable only if teams can deploy and operate services safely.
Before production extraction, establish:
- Delivery automation: Reproducible builds, automated tests, artifact storage, and environment promotion.
- Observability: Structured logs, metrics, distributed traces, and consistent request correlation.
- Reliability controls: Timeouts, bounded retries, health checks, and resource limits.
- Security: Workload identity, least-privilege permissions, managed secrets, and authenticated service communication.
- Ownership: A named team, service objectives, runbooks, and an escalation path.
GitHub Actions, GitLab CI/CD, and Jenkins can automate delivery. Terraform can provision infrastructure. Prometheus and Grafana support metrics, while managed observability platforms can reduce maintenance work.
Instrument both the monolith and extracted services using the official OpenTelemetry documentation. Cross-boundary visibility matters most during coexistence, when one customer request may traverse both architectures.
Kubernetes is optional. Amazon ECS, Google Cloud Run, Azure Container Apps, virtual machines, or existing application platforms may meet the requirements with less operational overhead. Choose based on workload behavior and team capability, not architectural fashion.
A step-by-step migration process
1. Establish scope, outcomes, and stopping rules
Create a short migration charter for the first capability. Include its business purpose, dependencies, owning team, baseline metrics, and expected benefit.
Set explicit go/no-go conditions:
- Functional behavior must match approved acceptance tests.
- Latency and error rates must remain within agreed objectives.
- Data reconciliation must pass before ownership changes.
- Rollback must be demonstrated in a production-like environment.
- Operational cost must fit the approved budget.
Schedule a review after stabilization. Continued extraction should require evidence, not merely an unfinished roadmap.
2. Choose a representative, manageable first service
Select a capability with a clear boundary, understandable data, and enough business relevance to test the approach.
Notification delivery, document generation, or a bounded catalog capability can be candidates, depending on the application. Avoid choosing solely because a module is easy: an isolated utility may teach little about domain ownership or transactional migration.
Conversely, beginning with the central checkout or financial ledger can concentrate too much risk before the team has proven its tooling.
3. Create a seam inside the monolith
Refactor callers to use an explicit interface instead of accessing implementation details directly. This is often called branch by abstraction.
Initially, the interface still invokes the monolith’s implementation. Later, it can route to the new service. This separates boundary creation from deployment change and makes rollback easier.
Add characterization tests around existing behavior, including inconvenient edge cases. Document which quirks must remain compatible and which changes require product approval.
4. Extract behind a strangler routing layer
With the strangler fig pattern, the new implementation gradually replaces selected functionality while the monolith continues serving everything else.
Routing can happen through an application interface, reverse proxy, API gateway, or event subscription. NGINX, Envoy, and managed API gateways are common options.
Microsoft’s Strangler Fig pattern guidance describes the incremental approach and its limitations.
Start with internal users or a controlled tenant cohort when feasible. Keep routing deterministic for stateful workflows so related operations do not unexpectedly alternate between implementations.
5. Transfer data ownership deliberately
Code extraction is not complete while both systems freely write the same tables. Define one authoritative writer for each business entity at every migration stage.
A common sequence is:
- Define the service-owned schema.
- Backfill existing records.
- Capture ongoing changes while backfill runs.
- Reconcile records and business invariants.
- Switch reads when freshness and correctness are acceptable.
- Transfer write ownership through a controlled cutover.
- Remove old access paths after stabilization.
Coordinate the backfill and change stream around a consistent snapshot or recorded log position. Otherwise, records can be missed or applied out of order.
Debezium can capture database changes, while Kafka can transport them. However, change data capture reproduces database changes; it does not automatically create meaningful domain events.
6. Test contracts and failure behavior
Unit tests are necessary but insufficient. Add:
- Consumer-driven contract tests: Pact can verify expectations between service consumers and providers.
- Integration tests: Testcontainers can provide disposable databases and brokers.
- Workflow tests: Exercise critical business journeys across old and new components.
- Failure tests: Simulate timeouts, unavailable dependencies, duplicate messages, and delayed events.
- Compatibility tests: Verify that old and new application versions can coexist during deployment.
Keep a focused end-to-end suite rather than relying exclusively on a large, slow test environment.
7. Roll out gradually and prove rollback
Canary releases expose a limited traffic segment to the new service. Compare errors, latency, resource consumption, and business outcomes against the previous implementation.
Shadow traffic can help compare read responses, but do not replay production writes into an active system without preventing side effects.
Rollback has two parts: restoring traffic routing and preserving data correctness. Once the new service accepts writes, routing back is safe only if the old implementation can access or reconstruct those changes.
Use backward-compatible schema changes and delay destructive migrations. For some incidents, a forward fix is safer than returning to a stale database.
8. Stabilize and remove the old path
After acceptance criteria hold under representative load, remove obsolete code, database permissions, jobs, and routing rules.
Update runbooks, dependency maps, ownership records, and cost allocation. Temporary integration layers become permanent complexity unless their removal is explicitly planned.
Handle distributed data without recreating the monolith
The hardest architectural change is often losing a single database transaction across an entire workflow.
Choose consistency per business invariant
Classify requirements explicitly:
- Strong local consistency: Rules such as balanced ledger entries should generally remain within an appropriate transactional boundary.
- Eventual consistency: Search indexes, analytics, and many read models can tolerate delayed updates.
- Coordinated workflows: Order fulfillment may need several local transactions with explicit recovery behavior.
For multi-service workflows, a saga coordinates local transactions and compensating actions. Compensation is not always an exact undo: refunding a payment is a new business action.
Use a transactional outbox to persist state changes and outgoing event records in the same local database transaction. A relay publishes those records afterward. AWS’s transactional outbox guidance explains the dual-write problem this pattern addresses.
Consumers still need idempotency because deliveries can repeat. Define event identifiers, deduplication behavior, ordering requirements, and schema evolution rules.
Choose synchronous and asynchronous communication selectively
HTTP or gRPC works well when a caller needs an immediate answer. Messaging through Kafka, RabbitMQ, or Amazon SQS can decouple background work and absorb bursts.
Asynchronous communication introduces lag, duplicates, poison messages, and replay considerations. Synchronous communication introduces runtime dependency chains.
Avoid long call chains on customer-facing paths. Propagate deadlines and constrain retries so a failing dependency does not trigger a retry storm.
Common migration mistakes
- Creating a distributed monolith: Services require coordinated releases and share database internals. Fix ownership and contracts before adding more services.
- Splitting into tiny services immediately: Network and operational overhead exceed the benefit. Start with cohesive capabilities and refine using evidence.
- Ignoring non-request workloads: Reports, batch jobs, and database triggers continue depending on extracted tables. Include them in discovery.
- Using unrestricted dual writes: Partial failure leaves databases inconsistent. Prefer controlled ownership and durable propagation patterns.
- Centralizing every deployment decision: Services are nominally independent, but teams remain blocked by a central queue. Offer reusable platform defaults with clear autonomy.
- Underfunding coexistence: The migration temporarily runs duplicate infrastructure and integration mechanisms. Budget for that transition explicitly.
Measure value before expanding
Review each extraction against its original constraint. Has release coordination decreased? Can the capability scale independently? Are incidents easier to contain, or has operational work simply moved?
Track both benefits and costs: delivery performance, service objectives, cloud spending, on-call load, and cross-team dependencies. Normalize comparisons for workload and scope changes.
Proceed when the extracted service creates durable value and has sustainable ownership. Pause when boundary disputes, reliability regressions, or support costs dominate. Internal modularization may be the better next investment.
For related platform planning guides, browse more Migration topics.
Frequently asked questions
How long does a monolith to microservices migration take?
There is no reliable universal timeline. A bounded extraction may fit into a team’s delivery planning horizon, while broader modernization can span multiple planning cycles. Estimate discovery, refactoring, data transfer, coexistence, and decommissioning separately.
Should every monolith become microservices?
No. A modular monolith often suits a small team or a tightly coupled domain. Migration is justified when independent deployment, scaling, ownership, or isolation delivers benefits that outweigh distributed-systems complexity.
Can microservices share a database during migration?
Temporarily, yes, but ownership must remain explicit. Restrict writes to a single owner and track shared access as migration debt. Separate schemas and permissions can be intermediate steps; direct cross-service table access should not become the permanent integration contract.
How do you avoid downtime during migration?
Use compatible schemas, incremental synchronization, deterministic routing, and gradual rollout. Rehearse write-ownership transfer and recovery. Some systems still need a controlled write pause to preserve correctness; avoiding every pause should not take priority over protecting business data.
Ask the community and get answers from practitioners.