We split a monolith into services and p95 latency dropped about 40 percent. That number is the least interesting thing that happened: the real win was failure isolation, and the real costs are the ones nobody puts in the conference talk.
What stopped happening
One bad deploy used to take the whole product down. After the split it took one service down, and the other six kept serving. That is the actual reason to do this. Latency was a side effect.
The honest cost
Local development got worse. Running the full product on a laptop went from one command to a compose file nobody fully understood.
Debugging moved from a stack trace to correlating three logs across two services. Distributed tracing stopped being nice to have and became the thing you install before you split, not after.
The trap
If your monolith is slow because of one endpoint doing an N+1 query, splitting it will not help. You will have the same slow query, in a smaller box, with a network hop in front of it.
Takeaways
- The main benefit of splitting is failure isolation, plus clearer team ownership.
- Expect worse local development and harder debugging.
- Set up distributed tracing before the split, not after.
- If you are splitting for speed, profile first. The answer is usually a missing index.
Building something like this?
I'm Ahmed Mamdouh, a senior full-stack & AI engineer. I reply within one working day.
Watch queue depth trend, not the current number
Monitor queue depth over time from day one; the current number says little, but the slope tells you whether you have twenty minutes or two.
Scaling 100k WebSocket connections: the reconnect storm
At 100k+ concurrent sockets, the hard part is not the count but the reconnect storm; jittered backoff, load shedding and resumable sessions fix it.
npm v12 blocks install scripts by default
npm v12 no longer runs preinstall, install or postinstall scripts unless you approve them, closing a common supply chain attack path.