Terraform state, four teams, one repository
Splitting a monolithic state file without a maintenance window, and the two mistakes that cost us a weekend.
I inherited a single state file with 4,200 resources. Every plan took eleven minutes and every apply became a negotiation between four teams.
We moved in slices by using moved blocks and state mv in the same change, then froze the source of truth for exactly one hour per slice.
The first mistake was assuming import would preserve lifecycle ignore_changes rules. It does not. The second was running the migration on a Friday.
I split by ownership boundary, not resource type. Networking, platform, application and data each got a dedicated state, with shared modules versioned explicitly. That reduced plan noise immediately because each team stopped reviewing unrelated drift from other domains.
The migration plan for each slice had four fixed steps: lock both states, declare moved blocks, run state mv with a reviewed mapping file, and run plan in both source and destination before any apply. We treated any unknown change as a stop signal, even when the diff looked harmless.
Remote-state references were the most fragile edge. A few modules consumed outputs indirectly through wrapper locals, and those paths broke silently after the move. We added a temporary compatibility output layer and removed it only after downstream repos had pinned updated module versions.
State backup discipline mattered. Before each slice, I stored encrypted state snapshots and the exact mapping manifest used for moves. During one rollback, that manifest reduced recovery time from an estimated hour to under ten minutes.
By the end, median plan time dropped from eleven minutes to under three, and change review quality improved because plans were small enough to reason about. The biggest lesson was procedural: migrations like this need weekday support windows and named owners for every dependency chain.