01The problem
Every IPTV device on the ALTO platform keeps a live connection for commands, status and notifications. At 100,000+ devices, the backend has to hold every socket open, route messages to the right device and survive spikes without dropping events.
On top of that, the original backend was a monolith: one slow path could drag the whole platform's latency up.
02Approach
I designed the real-time device management system on NestJS with Socket.io for the device connections and MongoDB for device state.
Notifications run through a WebSocket and queue-based event pipeline that delivers millions of events a day with at-least-once semantics, so a slow consumer delays a message instead of losing it.
We split the monolith into NestJS services on Docker and Kubernetes. Isolating failure domains cut p95 latency by about 40%.
03Architecture
Devices connect to NestJS Socket.io gateways; events flow through queues to NestJS services backed by MongoDB.
A schema-first GraphQL API is the single layer for web, mobile and partner integrations.
Everything ships through Jenkins CI/CD to Kubernetes on AWS EKS, with monitoring and incident response owned by the team I led (7+ engineers).
04Results
100,000+ concurrent device connections, millions of events a day, and p95 latency down ~40% after the services split. The same device telemetry later fed the ALTO AI agent and MCP server.
FAQ
What is ALTO?
How do you keep that many WebSocket connections reliable?
Can you build a real-time system like this for me?
Need something like this?
I'm Ahmed Mamdouh — 10+ years building systems like ALTO device platform. Scoped in one call.