Story / N°08 · DevOps

The story of Titan Cloud Migration.

Two hundred and forty services, one continent, zero downtime.

08 · DEVOPS2023 · 32 weeks
01 — Six-hour deploys, quarterly outages

The retailer's estate was 240 services on aging on-prem VMware. Deploys took six hours. Outages arrived on a quarterly rhythm. The board had approved a cloud migration in principle three years running, and every attempt had stalled inside its first quarter.

02 — Codify before you move

We refused to lift anything until it was codified. Terraform first, then a service catalog, then a pipeline that could stamp out an environment from scratch. When the migration actually began, every service moved via the same three-step recipe.

03 — Traffic mirroring, not big-bang

Istio's traffic mirroring became the workhorse. Every service ran shadow traffic in GCP for a full week before a single real request routed to it. The team caught 61 latent bugs in mirror mode; the customer never saw one of them.

04 — Forty engineers, thirty-two weeks

In parallel with the migration, we ran a training program for 40 internal engineers so the platform was owned in-house on day one after we left. Migration completed with zero customer-visible downtime. Deploys are now nine minutes. Total infrastructure spend is down 43% year-over-year.

The cloud program used to be a punchline. Now it's the platform.
CTO, Fortune 500 Retail
— Epilogue

The retailer has since retired the last of its on-prem VMware footprint and is running two additional GCP regions on the same platform.