Karla Burnett
Security Engineer
Lorikeet
Deliberate Holes Only
Deliberate Holes Only
Making an AI agent vulnerable to exactly one attack is harder than making it completely secure, let alone doing it six different times.
In this talk, I'll describe the process of building an agentic AI capture the flag challenge (https://owngoal.lorikeetcx.ai), in which each level has a more sophisticated agent than the last, always with one hole left for a player to find. Hundreds of people played the game, but only around 30 were able to complete it successfully.
I'll cover:
- choosing each level's vulnerability, from trivial prompt injection through to subtle social engineering, and using evals to prove nothing else got through;
- where the built-in defences from frontier labs helped, and where they got in the way;
- rewriting agent responses on the fly to hint solutions to players, and the security tradeoff that creates;
- shipping in a real environment: last minute model deprecations and a full retheme from marketing, along with the evals that made both survivable.
Attendees will leave knowing which types of vulnerabilities to watch out for in their own public-facing agents, and how to build evals to make sure there are no unexpected holes in them.
Karla Burnett
I'm a security- and infrastructure-focused software engineer at Lorikeet, an AI customer support automation platform for regulated industries. At Lorikeet, I've built production AI systems including our RAG stack, migrated our ticket processing pipeline to Temporal, built the public-facing CTF I'll be speaking about, and cut our infrastructure costs by 70% in the space of 2.5 months.
Before Lorikeet, I spent eight years at Stripe across security and product engineering. There, I rewrote Stripe's web authentication stack, helped build the first version of Stripe Sigma, cut account-takeover losses in half, built systems for securely executing arbitrary code against production, and, at various points, phished the entire company.