Yulia Kuchina
Staff AI Engineer
Software at Scale
How to Change an LLM System Without Guessing
How to Change an LLM System Without Guessing
We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide their regressions.
This talk is how we went from operating on faith to changing the system on evidence — a systems story, not a model story. The layer, built in order: prompt fingerprinting so every output traces to the exact prompt and model; provenance on every document record; failure visibility so classification failures surface as data, not silence; and an evaluation harness wired into CI, scoring every change against a labelled baseline — report-only today, becoming a blocking gate once the corpus supports a threshold we trust. The corpus is the honest bottleneck — a starving eval set can’t gate anything — so we’re building a pipeline that mints eval fixtures from real production failures without PII ever entering git: labels in version control, documents in an erasable store inside the production boundary.
The trade-offs, honestly: why we grounded extraction in document citations the model has to resolve — fabricated evidence fails to resolve instead of producing a false highlight — why a lawyer’s override is becoming first-class state the system can re-check, and the decisions I’d make differently. I’ll show a regression the harness surfaced before it shipped, and a defect that sailed past every automated check: a legally-wrong output that was mechanically correct, in a dimension none of our instruments measured — where human judgment takes over.
The takeaway: the reliability of an AI feature lives in the harness around the model, not the model. Prompts, parsers, evals, telemetry, and release gates are one system — once you can measure it, you stop guessing.
Yulia Kuchina
Yulia Kuchina is a Staff AI Engineer at Software@Scale, currently working with Commonwealth Bank on rebuilding enterprise banking software with AI coding agents and developing ways to evaluate the code they generate.
She specialises in making LLM systems measurable and safe to change, with hands-on experience across evaluation, provenance, observability, citation grounding, confidence routing and production failure analysis. Previously, at LEAP legal software, she owned a production AI document pipeline and built evaluation and regression controls for assessing changes before release.
Her background spans 12 years across frontend, full-stack and AI engineering in banking, legal technology, government and e-commerce. She is particularly interested in replacing impressive demos and intuitive prompt changes with evidence: production-derived test cases, versioned baselines and metrics that reveal whether a system has actually improved.