Online · Pacific Time (PT)

Agentic AI Frontier Seminar

A seminar series on Agentic AI: models, tools, memory, multi-agent systems, online learning, and safety, featuring leading researchers and industry experts.

Join us! Register here to get seminar updates.

Recordings (with speaker consent) will be posted on our YouTube channel.

Incoming Seminar

Online · 2026-09-02 · 08:00–09:00 PT

Talk Title: Toward Reliable Deployment of LLM Agents: From Uncertainty Quantification to Progress Advantage

Associate Professor · Sharon Li · University of Wisconsin-Madison

LLM agents are increasingly deployed on long-horizon tasks with tool use, irreversible actions, and unpredictable feedback. Yet we have few principled ways to tell, mid-episode, whether an agent is on track or quietly failing. Most uncertainty quantification (UQ) research still centers on single-turn QA, a poor match for interactive agents. In this talk, I'll present a general formulation of agent UQ and the challenges unique to agentic settings, from choosing uncertainty estimators to modeling how uncertainty evolves over an interaction. I'll then show that a powerful answer has been hiding in plain sight: RL post-training already yields an implicit step-level signal, the progress advantage, which recovers the optimal advantage function with no annotation or reward-model training. Across test-time scaling, UQ, and failure attribution, this free byproduct beats confidence baselines and even dedicated trained reward models.

Bio: Sharon Li is an Associate Professor in the Department of Computer Sciences at the University of Wisconsin-Madison. Her research focuses on algorithmic and theoretical foundations of reliable machine learning, addressing challenges in both model development and deployment in the open world. Previously, she was a postdoc researcher in the Computer Science department at Stanford University. She completed her Ph.D. from Cornell University, advised by John E. Hopcroft. She has served as the Program Chair for ICML 2026. She was the recipient of Alfred P. Sloan Fellowship (2025), NSF CAREER Award (2023), MIT Innovators Under 35 Award (2023), AFOSR Young Investigator Award (2022), Forbes 30under30 in Science (2020), and multiple faculty research awards from Google, Meta, and Amazon. She was named the “Innovator of the Year 2023” by MIT Technology Review. Her work has won the Outstanding Paper Award at NeurIPS 2022 and ICLR 2022.

Organizing Committee

Photo of Ming Jin

Ming Jin

Virginia Tech

He is an assistant professor in the Bradley Department of Electrical and Computer Engineering at Virginia Tech. He works on trustworthy AI, safe reinforcement learning, foundation models, with applications for cybersecurity, power systems, recommender systems, and CPS.

Photo of Shangding Gu

Shangding Gu

Shanghai Jiao Tong University

He is an associate professor in the School of Computer Science at Shanghai Jiao Tong University. His research focuses on reinforcement learning, planning, and AI safety, with applications in foundation models (e.g., large language models and multimodal models), robotics, and semiconductor manufacturing.

Photo of Yali Du

Yali Du

KCL

She is an associate professor in AI at King’s College London. She works on reinforcement learning and multi-agent cooperation, with topics such as generalization, zero-shot coordination, evaluation of human and AI players, and social agency (e.g., human-involved learning, safety, and ethics).

Photo of Lifu Huang

Lifu Huang

UC Davis

He is an Associate Professor in the Computer Science Department at UC Davis. His research centers on Natural Language Processing, Machine Learning, and Artificial Intelligence. His current research focuses on vision-language models, agentic AI, and the robustness of RL-based post-tuning.

Photo of Chenguang Wang

Chenguang Wang

UC Santa Cruz

He is an assistant professor in the Department of Computer Science and Engineering at UC Santa Cruz, and a research advisor at Scale AI. His research focuses on natural language processing, machine learning, and security, including AI agents, foundation model evaluation, and the safety and security of large language systems.