Harry Nguyen
AI engineer
insightfactory.ai
Running 100K+ Concurrent Agents: It's Never the Prompt
Running 100K+ Concurrent Agents: It's Never the Prompt
Your Agents Aren't Failing at Prompting
Every agent demo works. Then you run ten thousand at once, and the failures have nothing to do with the model: dead tuples piling up, a sync HTTP client freezing the event loop, requests hanging forever with no total timeout.
This talk is a look inside how we build and run agents at massive scale in production. It covers what broke, what fixed it, and the lessons that stuck, from why prefetch is really a distributed semaphore, to why layered timeouts beat one big number, to how Postgres as a queue can drive autoscaling at massive scale. It also shows how to get the most out of Postgres and its cache before bringing in Redis.
Harry Nguyen
I'm an AI Engineer at insightfactory.ai, where I build multi-agent systems at massive scale. I studied Computer Science at the University of Adelaide, and previously worked as a Machine Learning Engineer at the Australian Institute of Machine Learning, where my research focused on adversarial attacks and backdoor attacks against state-of-the-art object detectors, published at top-tier security conferences.