Dean Soste

Dean Soste

Senior Machine Learning Engineer

Canva

Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals

See all speakers

Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals

Production AI problems are rarely observable through deterministic CI checks. They surface as user feedback, bad LLM-as-a-Judge scores, alerts, or even a vague sense that something is wrong. The evidence that explains them is scattered across traces, prompts, content, documentation, support tickets and logs.

At Canva, we treated that evidence as a graph and built an internal MCP layer that lets agents traverse it on demand, with every claim retaining a path back to its source.

Our teams use these tools to find production bugs; iterate on prompts and content; correlate text across Java code, Langfuse prompts and Contentful entries; run agentic end-to-end tests; and find examples of specific behaviours buried in production traces. 70 people have made over 200,000 tool calls, and one $8.87 session uncovered three reproducible bugs in production.

Rather than building a graph database or eval agent, we focused on tools that mirror the process of a good human investigator: treating the initial signal as a lead rather than a concrete fact, searching across systems and providers, retrieving detail and domain knowledge only when needed, tracing findings to their sources, exposing gaps, and identifying where to look next.

Two constraints shared by LLM systems become critical in longer-running agents: every problem has a context floor, but every token fights for attention. Too little context leaves an agent guessing; too much dilutes attention, raises cost and accumulates across the task. The balance is what we call Minimally Sufficient Context: the smallest set of information that still contains everything an agent needs to reach a justified conclusion and a human needs to verify it.

Usage has also exposed clear limits: agents miss inaccessible evidence, conclude too early and struggle to judge real-world impact. This talk turns those lessons into a practical build order for agentic evals: make evidence navigable, test tools with people on real problems, observe how they are used, and automate only what proves useful.

Dean Soste

Dean Soste is a Senior Machine Learning Engineer at Canva, specialising in AI evaluation and the infrastructure needed to improve production AI systems. He has integrated production telemetry into Canva’s core help-domain AI products: the Help Assistant and Omni-agent, its support-ticket responder. His work supports a help ecosystem serving Canva’s 250 million monthly active users and handling roughly 180,000 support tickets each month.