Online · Pacific Time (PT)

Agentic AI Frontier Seminar

A seminar series on Agentic AI: models, tools, memory, multi-agent systems, online learning, and safety, featuring leading researchers and industry experts.

Join us! Register here to get seminar updates.

Recordings (with speaker consent) will be posted on our YouTube channel.

Incoming Seminar

Online · 2026-09-16 · 08:00–09:00 PT

Talk Title: Building Reliable Evaluations for Frontier Agents

Research Scientist · Daniel Zhang · Scale AI

As frontier agents tackle increasingly complex, real-world work, the grading systems that evaluate their performance are now doing double duty-serving as both reinforcement learning rewards and mechanisms for test-time selection. This talk explores how Scale AI builds robust evaluations for cutting-edge coding agents and beyond, featuring insights from building leading leaderboards such as SWE-Bench Pro, SWE Atlas, Terminal-Bench, and MCP Atlas. It also examines the critical challenges of designing reliable verifiers, comparing programmatic versus rubric-based grading, identifying and mitigating reward hacking, and architecting agentic verifiers capable of auditing long-horizon tasks.

Bio: Daniel Zhang is currently Head of Agents Research at Scale AI, where his team focuses on frontier agent research in Coding, Computer Use, Long-horizon Tool Use, and Safety. Previously, he served as Post-training Lead at Amazon AGI, where he worked on reasoning and coding agents. He holds a Ph.D. in Computer Science from the University of Notre Dame and an M.S. in Information Security from Purdue University.

Organizing Committee

Photo of Ming Jin

Ming Jin

Virginia Tech

He is an assistant professor in the Bradley Department of Electrical and Computer Engineering at Virginia Tech. He works on trustworthy AI, safe reinforcement learning, foundation models, with applications for cybersecurity, power systems, recommender systems, and CPS.

Photo of Shangding Gu

Shangding Gu

Shanghai Jiao Tong University

He is an associate professor in the School of Computer Science at Shanghai Jiao Tong University. His research focuses on reinforcement learning, planning, and AI safety, with applications in foundation models (e.g., large language models and multimodal models), robotics, and semiconductor manufacturing.

Photo of Yali Du

Yali Du

KCL

She is an associate professor in AI at King’s College London. She works on reinforcement learning and multi-agent cooperation, with topics such as generalization, zero-shot coordination, evaluation of human and AI players, and social agency (e.g., human-involved learning, safety, and ethics).

Photo of Lifu Huang

Lifu Huang

UC Davis

He is an Associate Professor in the Computer Science Department at UC Davis. His research centers on Natural Language Processing, Machine Learning, and Artificial Intelligence. His current research focuses on vision-language models, agentic AI, and the robustness of RL-based post-tuning.

Photo of Chenguang Wang

Chenguang Wang

UC Santa Cruz

He is an assistant professor in the Department of Computer Science and Engineering at UC Santa Cruz, and a research advisor at Scale AI. His research focuses on natural language processing, machine learning, and security, including AI agents, foundation model evaluation, and the safety and security of large language systems.