Sequencing migration waves around blast radius
The order you move workloads in determines whether a failure is an incident or an outage. A defensible wave plan starts with the lowest-risk target.
Sequencing migration waves around blast radius
The order you move workloads in determines whether a failure is an incident or an outage. A wave that takes down one internal tool is an incident. The same wave taking down customer transactions during business hours is an outage, and it will be handled differently by everyone involved.
Sequencing is therefore a risk decision, not a project-management convenience.
Sequence by blast radius, not by ease
The instinct is to start with the easy win: a stateless service with no database dependencies, moved in a morning. It builds confidence and it teaches the team.
That is worth something. It is not worth it if it spends your rehearsal capacity on the workload that cannot hurt you.
A defensible order:
- Non-customer-facing, low criticality. Internal tooling, staging environments, a reporting replica.
- Customer-facing but low consequence. A marketing site, a documentation portal, an internal read-only dashboard.
- Write-capable, reversible. A system with a tested rollback and a data model you have already copied.
- Transaction processing. Only after two or three waves have run cleanly through a full business cycle, including month-end.
Most migrations fail at wave four, and they fail because waves one to three consumed the goodwill, the change window and the attention that wave four needed.
Count dependencies before choosing a wave
A workload is not independent because its team says so. In practice:
- a reporting job reading a production database
- an authentication provider every internal system depends on
- a file store holding documents other systems fetch by path
- a scheduled export consumed by a partner
- a message queue whose consumers assume a single producer
None of these appear on an architecture diagram that shows services. They appear in network flows, cron schedules and access logs.
The most reliable discovery method is access logging over a representative period, not interviews. Ask what should happen, then check what the logs say happens.
Rehearse the cutover, do not plan it
A cutover that has never been executed is a hypothesis.
Rehearse against a production-shaped environment: production data volumes, production network conditions, and the same team. Time every step. The rehearsal exists to produce two artefacts — a realistic duration, and the list of steps that only fail when you run them.
The steps that fail during rehearsal are the valuable output. Expect several, and expect them to be the batch jobs.
Test the rollback, not just the cutover
Every migration plan has a rollback step. Few have tested it.
Testing matters more than teams expect, because rollback paths degrade over time: a script that referenced a table which has since been renamed, credentials that expired, a colleague who wrote it having left.
Make the rollback a rehearsed procedure with a named owner and a time estimate. Know how long the decision takes, not just the action. A rollback that takes four hours is not a rollback; it is a longer outage.
Align waves to business calendars
Some weeks are worse than others. Avoid:
- month-end and quarter-end close
- peak trading or production periods
- the weeks a compliance report is due
- the week your most experienced operator is on leave
Also avoid a wave in the fortnight after a previous one. A migration produces a stabilisation period; scheduling another wave into it is how small issues compound into a difficult month.
Define the abort criteria in advance
For each wave, write down before cutover what causes you to stop:
- transaction error rate above a threshold
- reconciliation mismatch
- a named journey failing its synthetic check
- elapsed time beyond the rehearsed duration plus a margin
Deciding these during an incident means the decision gets made under time pressure by exhausted people. Deciding them in advance means you can hand the decision to whoever is authorised to make it, with the criteria already agreed.
Be honest about what success looks like
A wave is not successful when the workload is running in the new environment. It is successful when:
- the workload has run a full business cycle including any scheduled jobs
- reconciliation matches for the period
- support volume is at baseline
- the rollback path has been documented as still valid, or retired
Most organisations declare victory at cutover and discover in week three that a monthly job did not migrate. Build the definition before you start, and enforce it.
In this article
- cloud migration
- cutover
- risk
Working on something similar?
These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.
Start a conversationRelated reading
Continue from here
Articles connected to the same delivery problems.
Which four metrics are worth alerting on
Have a related problem in front of you?
Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.