Hire an AI agent and MCP server developer
In my current role I built an AI agent on the Anthropic and OpenAI APIs and the MCP server it uses to reach internal systems. Both run in production. Most of my 10+ years went into backends, and that turns out to be most of the work in an agent: state, partial failures, permissions and knowing when to stop.
In production
Anthropic and OpenAI APIs with LangChain: planning, tool calls, short-term memory, guardrails on every action.
Tool server for LLMs
Read tools answer questions. Action tools need a human to approve them. Destructive actions aren't exposed.
Retrieval with evals
Embeddings in Qdrant, and an eval harness that runs on every prompt or model change.
Fine-tuning
On S3, SageMaker and Bedrock, with evals to catch regressions.
What I build
An agent or AI feature usually takes 2 to 4 weeks to reach production.
An agent inside your product or ops
Tool calls into your APIs, short-term memory, and guardrails that check each action before it runs.
An MCP server for your systems
Scoped tools so Claude and other MCP clients can query your data and request actions without being able to break anything.
RAG over your docs and data
Structural chunking, metadata filters and a small eval set. In my experience retrieval fixes beat model upgrades.
Evals, usage and cost tracking
So you can see when an answer gets worse or a feature gets expensive, before your users do.
The backend around it
Queues, retries, auth and deploys on AWS. Agents fail like distributed systems, so they need the same engineering.
Streaming UIs
Chat and assistant interfaces in Next.js with the Vercel AI SDK, like Ask my CV on this site.
Case studies
The engineering behind the numbers above. Products under NDA are described without internals.
AI agent + MCP server
A production AI agent with tool use and guardrails, plus an MCP server that lets LLM clients query and act on internal systems under strict rules.
Read case studyReal-time device platform
A device management backend that keeps 100,000+ devices connected at once. The product is under NDA, so this page covers the engineering, not the product.
Read case studyHow I think about this work

What actually breaks when AI agents go to production
Production agents fail on state, partial failure and knowing when to stop. Those are distributed systems problems, and prompts won't fix them.

In RAG, retrieval is the product
Most bad RAG answers I've debugged were retrieval failures. Structural chunking, metadata filters and a 50-question eval set fixed more than model upgrades.

MCP went stateless, and that matters
The July MCP revision went stateless with OAuth 2.0 and OpenID Connect, so MCP servers can now run behind load balancers and serverless.
From first call to launch
- 01Day 0
Intro call
15 minutes on what you are building and by when.
- 02Day 2
Scope & quote
A written plan with milestones and a fixed price.
- 03Weekly
Build in the open
Weekly demos, a staging link and async updates.
- 04Final week
Launch & handover
Deploy, docs and a recorded walkthrough.
Questions clients ask
Have you shipped AI agents to production?
What does an MCP server do for my product?
Can you add RAG to an existing product?
Claude or OpenAI?
How do we work together?
Tell me what you're building
I'm Ahmed Mamdouh, a senior full-stack and AI engineer in Cairo, Egypt. I reply within one working day.
Looking for something else? See all the ways to work together or nestjs developer.