An endpoint got about 45 percent faster and I did not write a single clever line of code. I ran a profiler, and it disagreed with four experienced engineers.
A week of guessing
The team had been arguing about it for a week. One person was sure it was the ORM. Another wanted to move the whole thing to a queue. Someone suggested rewriting it in Go, because someone always does.
I ran a profiler for twenty minutes instead.
What the profile found
Three things came out, and none of them were what anyone guessed.
- A relationship loaded in a loop. One request became sixty one queries, the classic N+1.
- A WHERE clause on an unindexed column. It had been fine at ten thousand rows and was not fine at four million.
- A cache keyed per user. We had added it six months earlier. The hit rate was close to nothing, while we paid the write cost on every request.
An afternoon of fixes
Fixing all three took an afternoon. Eager loading, one index, delete the cache.
The actual lesson
The lesson is not "use a profiler". Everyone nods at that and goes back to guessing. The lesson is that the profile disagreed with four experienced engineers, and the profile was right. We would have spent a sprint rewriting the wrong thing.
Takeaways
- Profile before debating architecture changes.
- Look for N+1 queries, missing indexes and caches with poor hit rates.
- Queries that are fine on small tables can fall over as data grows.
- A cache that rarely hits is pure cost; deleting it is a valid fix.
- Measure first. It is boring and it keeps being correct.
Building something like this?
I'm Ahmed Mamdouh, a senior full-stack & AI engineer. I reply within one working day.
Node.js moves to one major release a year
From Node 27, Node.js ships one major a year and every release becomes LTS, ending the odd/even split most teams already ignored.
Mongo or Postgres is the wrong first question
Pick a database by your known access patterns; if you do not know them yet, pick Postgres, because it forgives wrong guesses best.
Caching mistakes I keep debugging
Before adding a cache, measure first, decide what invalidates it, set a TTL, and cache the expensive part rather than the whole result.