"Building a standalone AI agent that looks great in a local demo is relatively straightforward. The real headache starts when you connect multiple autonomous agents together, hook them up to enterprise data pipelines, and deploy them to real users. Non-deterministic outputs, unexpected model drift, and security vulnerabilities like prompt injection quickly turn a clean codebase into an operational nightmare. In this session, we'll step away from the hype and look at what actually happens when multi-agent architectures meet enterprise-scale workloads. Drawing from real-world engineering experiences, I'll walk through the specific architectural bottlenecks that crop up during orchestration and state management. We will look at practical, open-source safety and evaluation frameworks designed to continuously stress-test these systems before they hit production. Attendees will walk away with a clear blueprint for building robust, multi-layered guardrails that keep autonomous systems predictable and secure without tanking processing latency or driving up API costs."
Building a standalone AI agent that looks great in a local demo is relatively straightforward. The real headache starts when you connect multiple autonomous agents together, hook them up to enterprise data pipelines, and deploy them to real users. Non-deterministic outputs, unexpected model drift, and security vulnerabilities like prompt injection quickly turn a clean codebase into an operational nightmare. In this session, we'll step away from the hype and look at what actually happens when multi-agent architectures meet enterprise-scale workloads. Drawing from real-world engineering experiences, I'll walk through the specific architectural bottlenecks that crop up during orchestration and state management. We will look at practical, open-source safety and evaluation frameworks designed to continuously stress-test these systems before they hit production. Attendees will walk away with a clear blueprint for building robust, multi-layered guardrails that keep autonomous systems predictable and secure without tanking processing latency or driving up API costs.

Purva Chiniya is an Applied Scientist at Amazon Ads, where she works on LLM agents, safety inference, and large-scale machine learning systems. Her work spans multi-agent orchestration for advertiser decisioning, low-latency inference systems, and alignment methods for improving the safety and reliability of large language models. Previously, she interned at Amazon and Sentient Foundation, where she built agentic search and reasoning systems, safety steering mechanisms, and fine-tuning pipelines for large models.Purva’s research has been accepted at venues including LREC, LLM4Code, CVPR, EMNLP, and INTERSPEECH, covering topics such as LLM safety, secure code generation, agentic search, multimodal learning, and hate speech detection. She holds an M.S. in Computer Science from the University of Maryland, College Park, and a B.Tech. in Electrical Engineering from IIT Roorkee.