The gap between an agent demo and an agent in production is not the prompt. It is what happens on the third tool call, when the API returns a 500 and the agent has already written half a row to your database.
Demos are one happy path
Demos are one happy path. Production is partial failure, retries that must not double-charge anyone, and a clear answer to the question of when the agent stops and asks a person.
I keep seeing agent architecture diagrams. I rarely see anyone talk about the three things that actually break.
State
An agent that runs for ninety seconds and calls six tools is a long-running transaction with no transaction manager. If it dies on step four, something has to know that steps one to three happened. Most demos keep this in memory and hope.
Partial failure
A tool call times out. Did it not run, or did it run and lose the response? Those need different handling, and the agent usually cannot tell them apart. So every side-effecting tool needs an idempotency key, the same way every queue consumer does.
Knowing when to stop
This is the hard one. An agent that asks a human too often is a slow form. An agent that never asks is a liability. The threshold is a product decision, not a prompt, and it belongs in code where you can test it.
Nothing new
None of this is new. It is retries, idempotency and backpressure, which backend engineers have been doing for twenty years. It is ordinary distributed systems work wearing a new hat. The model is the easy part.
Takeaways
- Persist agent progress so a crash mid-run can be recovered.
- Give every side-effecting tool an idempotency key.
- Treat timeouts as "unknown outcome", not "did not run".
- Put the human-escalation threshold in testable code, not in the prompt.
- Agent reliability is distributed systems engineering.
Building something like this?
I'm Ahmed Mamdouh, a senior full-stack & AI engineer. I reply within one working day.
MCP went stateless, and that matters
The July MCP revision went stateless with OAuth 2.0 and OpenID Connect, so MCP servers can now run behind load balancers and serverless.
Claude output is now watermarked
Claude output now carries watermarks and signed provenance by default; tell your users, and never treat a missing watermark as proof.
At-least-once delivery means at least once
Every queue consumer will eventually run twice, so give each one an idempotency key and check it before doing the work.