Migrating a Monolith to Microservices Without Downtime
The Strangler Fig pattern in practice: how we decomposed a 400k-line Rails monolith over 14 months while shipping features the whole time.
“We need to break up the monolith” is one of the most common (and most dangerous) architectural decisions a growing company can make.
Done wrong: 18 months of slow velocity, reliability regressions, teams blocked on each other.
Done right: gradual extraction that improves things incrementally.
Here’s how we did it right.
The Strangler Fig pattern
The metaphor: a strangler fig grows around an existing tree, eventually replacing it. You don’t tear down the old system — you build the new one around it until the old one withers.
In practice:
- Identify a bounded context to extract (we started with billing)
- Build a new service with its own database
- Route traffic to the new service via an API gateway
- Keep the monolith in sync during the transition (dual-write)
- Cut over and remove the old code
The monolith never goes down. Customers don’t notice.
What we extracted and why
We prioritised by: team pain × blast radius if we got it wrong × independence from other domains.
Phase 1: Billing & Subscriptions → highest team pain, well-defined bounded context
Phase 2: Notifications → fire-and-forget, easy to make idempotent
Phase 3: Search → Elasticsearch already external, natural extraction
Phase 4: User Auth → highest risk, last
Key decisions
Async-first: Services communicate via events (Kafka) wherever possible. Synchronous RPC only for user-facing reads that need consistency.
Shared-nothing databases: Each service owns its data. No shared Postgres schema. This is the hard constraint that forces good domain boundaries.
Observability from day one: Distributed tracing (Jaeger), centralised logging (ELK), per-service SLOs.
14 months, zero downtime incidents caused by the migration. The monolith is now down to 160k lines — we kept the parts that work.