Skip to content
Hire me
01 / StartXP 0%0/9
Sep 8, 2026 · 1 min · by Ahmed Mamdouh

What actually breaks when AI agents go to production

#ai#agents#mcp#distributed-systems

The gap between an agent demo and an agent in production is not the prompt. It is what happens on the third tool call, when the API returns a 500 and the agent has already written half a row to your database.

Demos are one happy path

Demos are one happy path. Production is partial failure, retries that must not double-charge anyone, and a clear answer to the question of when the agent stops and asks a person.

I keep seeing agent architecture diagrams. I rarely see anyone talk about the three things that actually break.

State

An agent that runs for ninety seconds and calls six tools is a long-running transaction with no transaction manager. If it dies on step four, something has to know that steps one to three happened. Most demos keep this in memory and hope.

Partial failure

A tool call times out. Did it not run, or did it run and lose the response? Those need different handling, and the agent usually cannot tell them apart. So every side-effecting tool needs an idempotency key, the same way every queue consumer does.

Knowing when to stop

This is the hard one. An agent that asks a human too often is a slow form. An agent that never asks is a liability. The threshold is a product decision, not a prompt, and it belongs in code where you can test it.

Nothing new

None of this is new. It is retries, idempotency and backpressure, which backend engineers have been doing for twenty years. It is ordinary distributed systems work wearing a new hat. The model is the easy part.

Takeaways

  • Persist agent progress so a crash mid-run can be recovered.
  • Give every side-effecting tool an idempotency key.
  • Treat timeouts as "unknown outcome", not "did not run".
  • Put the human-escalation threshold in testable code, not in the prompt.
  • Agent reliability is distributed systems engineering.

Building something like this?

I'm Ahmed Mamdouh, a senior full-stack & AI engineer. I reply within one working day.

More notesSep 11, 2026

MCP went stateless, and that matters

The July MCP revision went stateless with OAuth 2.0 and OpenID Connect, so MCP servers can now run behind load balancers and serverless.

Sep 8, 2026

Claude output is now watermarked

Claude output now carries watermarks and signed provenance by default; tell your users, and never treat a missing watermark as proof.

Sep 1, 2026

At-least-once delivery means at least once

Every queue consumer will eventually run twice, so give each one an idempotency key and check it before doing the work.