← M&A Integration Suite Execution Plan · Artifact 35 · how to read this suite

Cloud Migration Wave Plan

Download Word

Everything Cumberland Valley runs must leave Cheatham Mutual's data center before September 30, 2024, when TS-05 exits. This plan sequences that evacuation into five waves, each with entry conditions, exit evidence and a rollback position — and it opens with the arithmetic that determines whether the plan is possible at all. Issued December 4, 2023, after landing zone acceptance.

A wave plan is a sequencing document, and the sequence is set by dependency and risk rather than by convenience or by which team is ready. Wave 1 is chosen because it proves the path, not because it delivers value. The last wave is gated by member identity resolution rather than by anything infrastructure controls. In between, the rollback window narrows with every wave — and knowing where it closes is the difference between a migration and a gamble.

Table of Contents

Part I — The Constraint
  1. Egress Arithmetic
  2. Wave Design Principles
Part II — The Waves
  1. Wave Schedule
  2. Entry and Exit Criteria
  3. Rollback Windows
Part III — Control
  1. Recovery Evidence Per Wave
  2. Risks
Part I — The Constraint

1. Egress Arithmetic

Before sequencing anything, the plan answers one question: does the data physically fit through the available connection inside the window? It is a calculation anyone can do on day one, and it is the single most common source of late-discovered schedule failure in a data center exit.

Total estate in Cheatham Mutual's data center  =  420 TB
   of which cold historical archive  =  340 TB  — claims history, images, correspondence
   of which operational  =  80 TB  — moves with its systems, wave by wave

Circuit  =  1 Gbps  →  sustained migration allocation 400 Mbps
   (the balance carries live production and TSA traffic — the circuit is not idle)

400 Mbps  →  50 MB/s  →  ~4.3 TB/day
340 TB  ÷  4.3 TB/day  =  78 days continuous
Seventy-eight days assumes the transfer never stops, never contends with production, never fails, and never re-transmits. None of those hold. With realistic contention and retry, the archive alone consumes five to six months of the migration window — leaving the operational estate, the waves, the testing and the cutover to share what remains of ten. The archive does not go over the wire.

1.1 The physical transfer decision

InputValueNote
Archive to move340 TBCold; no incremental change during transit
Usable capacity per appliance80 TBAfter encryption overhead
Appliances required5Run in parallel, not in series
Turnaround per appliance18 daysShip, ingest, verify, return
Physical transfer works here for one reason that will not always apply: the archive is cold. Nothing writes to fifteen-year-old claims images, so an appliance can be filled, shipped and ingested without a reconciliation problem — whatever was on it when it left is still current when it arrives. Operational data cannot move this way, because it changes while the box is in transit, which is why the 80 TB operational estate moves over the wire inside its waves and only the archive goes by appliance.
The point worth carrying out of this section: this calculation costs an afternoon and it is the difference between a plan and a hope. The failure mode is not doing it — it is doing it in month eight, when the transfer is visibly not keeping pace, the appliance lead time is six weeks, and the data center lease does not move. Everything about the wave schedule below is downstream of this arithmetic.
Part II — The Waves

2. Wave Design Principles

PrincipleReasoning
Wave 1 proves the path, not the valueThe first wave's purpose is to exercise the runbook, the tooling, the cutover process and the rollback — on something the business can afford to have go badly
Dependencies before dependantsIdentity and network before the workloads that need them. Shared services before consumers.
Risk ascends⚠ Each wave is harder than the last, so capability grows alongside difficulty rather than being tested by it
One wave in flight at a timeParallel waves share the same scarce people. Two half-finished waves is worse than one finished one.
Every wave has a rollback positionUntil it does not — and Section 5 states exactly where that line is
Exit is evidenced, not declaredIncluding a restore test. A workload that has never been recovered in its new home is not migrated, it is copied.
"Wave 1 proves the path, not the value" is the principle most often overridden, and overriding it is a false economy. There is always pressure to start with something that matters, because the TSA meter is running and a low-value first wave looks like wasted weeks. But the first wave is where the runbook is wrong, the tooling has a gap, the network rule was missed and the rollback has never been rehearsed. Discovering all of that on a system nobody depends on costs days. Discovering it on claims adjudication costs the program.

3. Wave Schedule

WaveContentsVolumeWindowGated by
W-0Archive bulk transfer — historical claims, images, correspondence340 TBDec 2023 – Mar 2024Landing zone accepted; appliances provisioned. ⭐ Runs in parallel with W-1 and W-2 — it needs shipping, not people.
W-1Non-production environments; utility servers; file shares12 TBJan 2024Landing zone acceptance gate (LZ-01..LZ-10)
W-2Reporting, document management, internal tools18 TBFeb – Mar 2024W-1 exit; monitoring proven in production
W-3Care management platform (preserved per AD-07); utilization management14 TBApr – May 2024W-2 exit. ⚠ First clinically significant workload.
W-4Integration layer; interfaces; enterprise data warehouse22 TBJun – Jul 2024W-3 exit; warehouse recovery gate satisfied
W-5Core administration and enrollment — the last thing out14 TBAug – Sep 2024⚠⚠ Member identity resolution complete. Not an infrastructure dependency.
Total420 TBDec 2023 – Sep 2024TS-05 data center exit
Wave 5 is the whole program compressed into one row, and its gate is not technical. Core administration cannot move until the target's members exist correctly in ACME's systems — which requires identity resolution, which requires the clerical review queue to be worked, which requires knowing how large that queue actually is. The last and most important wave is therefore paced by a number the Clean Team Protocol made unknowable until after closing, and no amount of infrastructure readiness moves it. Everything in W-1 through W-4 can be on schedule and W-5 can still be waiting.
W-0 is deliberately not numbered in sequence, because it does not behave like a wave. It consumes shipping and logistics rather than engineering effort, it has no cutover, and it runs concurrently with W-1 and W-2 without competing for the same people. Treating the archive as a wave would serialize it behind work it does not depend on and consume months the schedule does not have.

4. Entry and Exit Criteria

4.1 Entry — every wave

4.2 Exit — every wave

CriterionEvidence
Workloads operating in Azure at production loadPerformance measured against the pre-migration baseline, not against expectation
Interfaces reconciledControl totals balance for a full processing cycle
Policy compliance cleanZero non-compliant resources under the landing zone policy set
Cost attributedTagged and appearing correctly in showback
Restore tested in the new location⭐ Actual recovery performed, evidence retained. See Section 6.
Monitoring and alerting operatingAlerts fire in test; on-call has received and acted on one
Source decommissioned or scheduled⚠ The old system is off or has a dated shutdown. A wave that leaves the source running has not reduced anything.
The last row is where data center exits quietly fail. Workloads migrate successfully, the team moves to the next wave, and the original servers keep running — because nobody wants to be the person who turned something off, and because "leave it a few weeks just in case" is always reasonable. Repeat that five times and the data center is still full on the day the lease ends. Migration is not the objective; the empty data center is.

5. Rollback Windows

Rollback is a real option early and becomes a fiction later. Stating where the line falls — in advance, in writing — is what prevents a team from assuming a safety net that no longer exists.

WaveRollbackWindowWhy it narrows
W-1WideIndefiniteNon-production. The source is untouched and nothing depends on the target.
W-2WideDaysRead-mostly workloads; source can be restarted with minimal reconciliation
W-3Narrow72 hours⚠ Clinical data written in the new location must be reconciled back if reverted
W-4Narrow24 hoursInterfaces have been repointed; reverting means repointing every consumer
W-5None after cutover⚠⚠ Claims adjudicating in the new location cannot be un-adjudicated. Forward fix only.
Wave 5 has no rollback, and that is stated plainly here rather than discovered at three in the morning. Once claims are adjudicating on ACME's platform for the combined book, reverting would mean unwinding payments, reversing remittances and re-processing against a system whose data is now stale. The mitigation is not a better rollback plan — there is no such plan. It is a longer parallel run, a tighter go/no-go, and the willingness to not proceed. A cutover with no reverse gear must be preceded by a decision point where "no" is genuinely available, which is why the Day 1 go/no-go pattern is repeated for this wave.
Part III — Control

6. Recovery Evidence Per Wave

Every wave exit includes a tested restore of at least one workload in its new location. This is not a formality carried over from the landing zone gate — it tests something different each time.

WaveTier of workloadsWhat the restore test proves at this point
W-1Tier 3 — deferrableThe backup configuration works at all, and someone knows how to run a restore
W-2Tier 3 / Tier 2Restore works at larger data volumes, within the stated RTO for the tier
W-3Tier 2 — importantClinical data restores with referential integrity intact, not merely with files present
W-4Tier 2, incl. the data warehouseThe new-build warehouse recovers — it inherited no regime, so this is its first proof
W-5Tier 0 — criticalRecovery inside the Tier 0 RTO of four hours, under production conditions
The tests get harder deliberately, in the same way the waves do. A first restore proves the mechanism exists. A restore at Wave 5 proves an organization that had never operated a cloud platform can recover a claims system inside four hours — which is a claim about people and process, not about configuration. Testing that capability for the first time on the last and most critical workload would be the same error as putting core administration in Wave 1.

⚠ The Wave 4 test carries particular weight. The enterprise data warehouse is the program's only genuinely new system, it inherited no backup regime and no restore history, and its retention obligation runs to seven years against a compliance requirement rather than an operational one. Its recovery gate conditions are set out in the Landing Zone Design and must be satisfied before the wave closes, not after.

7. Risks

RefRiskResponse
WR-01Identity resolution runs long, delaying W-5The dominant schedule risk on the program. No infrastructure mitigation exists. Tracked through the clerical review queue, not through migration status.
WR-02Appliance turnaround slower than plannedFive appliances run in parallel with float; archive transfer starts first and finishes months before it is needed
WR-03Source systems not decommissioned after their waveDated shutdown required at wave exit; tracked to the data center lease date rather than to team comfort
WR-04Performance in Azure differs from on-premises baselineMeasured at every wave exit against the pre-migration baseline. Rehost means "same behavior," which is testable.
WR-05Waves overlap because of schedule pressureOne wave in flight is a rule. Overlap consumes the same people twice and produces two incomplete waves.
WR-06MSP step-down boundary falls mid-waveStep-down boundaries aligned to wave boundaries, not to calendar months
WR-01 is worth stating separately because it is not a cloud risk and cannot be managed as one. Every other row here responds to something the migration team controls — bandwidth, sequencing, testing, staffing. The dominant risk to this plan is the size of a clerical review queue produced by data quality nobody was permitted to measure before closing. The migration can be flawlessly executed and still miss the data center exit, because the last wave is waiting on people reading member records one at a time. That is the shape of this program: the technical work is not the hard part, and the hard part is not technical.

Related artifacts: 20 — Application Disposition Matrix · 22 — TSA Schedule & Exit Plan (TS-05) · 23 — Data Migration & EMPI Strategy · 24 — Cloud Migration Strategy · 25 — Integration Architecture · 28 — Risk Register · 34 — Cloud Landing Zone Design · 39 — Interface Build & Cutover Log