Skip to content
Hire me
01 / StartXP 0%0/9
Sep 4, 2026 · 1 min · by Ahmed Mamdouh

Caching mistakes I keep debugging

#caching#redis#performance#system-design

Before you add a cache, answer one question: what invalidates it? Every caching bug I have debugged came down to skipping that question or to one of three related sins.

The question to answer first

If the answer to "what invalidates it?" takes longer than thirty seconds, or ends with "we will figure that out later", you are not adding a cache. You are adding a second source of truth that disagrees with the first one at an unpredictable time.

Half the caching bugs I have debugged were not about the cache at all. They were about nobody deciding, on the day it shipped, when the data was allowed to go stale.

Sin one: caching before measuring

The endpoint was slow because of an N+1 query. The cache hid it until traffic doubled. A cache in front of an unexplained slow path does not fix it, it postpones it.

Sin two: no TTL strategy

"We will invalidate manually" is a promise your future self will break. Even if you have explicit invalidation, a TTL is the safety net for the case you forgot.

Sin three: caching the wrong thing

Caching the computed result instead of the expensive part. Then one field changes and you rebuild everything. Cache the expensive piece, and assemble the cheap parts around it.

Takeaways

  • Decide what invalidates a cache on the day you add it.
  • Measure first; do not let a cache hide a slow query.
  • Always set a TTL, even if it is long.
  • Cache the expensive part, not the whole computed result.
  • Cache late, cache small. Boring rules, quiet pager.

Building something like this?

I'm Ahmed Mamdouh, a senior full-stack & AI engineer. I reply within one working day.

More notesSep 13, 2026

Watch queue depth trend, not the current number

Monitor queue depth over time from day one; the current number says little, but the slope tells you whether you have twenty minutes or two.

Sep 12, 2026

Scaling 100k WebSocket connections: the reconnect storm

At 100k+ concurrent sockets, the hard part is not the count but the reconnect storm; jittered backoff, load shedding and resumable sessions fix it.

Sep 9, 2026

How an endpoint got 45 percent faster with no clever code

Twenty minutes with a profiler beat a week of guessing: an N+1, a missing index and a useless cache made an endpoint 45 percent faster.