Talk Title: On the Interpretability and Self-Evolution of LLM Agents
Abstract: As Large Language Model (LLM) agents transition from static text generators to autonomous decision-makers, two critical challenges emerge: understanding their internal decision-making mechanics and scaling their capabilities without costly human supervision. This talk explores the core architecture of LLM agents through two complementary perspectives: mechanistic interpretability at the model-environment boundary and scalable self-evolution via co-training. In the first part, we examine the micro-mechanics of an agent loop through a minimalist study pairing a 1-layer, 1-head Transformer with an action-executing harness on Sudoku. We show how a lightweight model ($0.87\text{M}$ parameters) trained on easy puzzles generalizes to solve $99.8\%$ of extreme puzzles. Through circuit analysis, we trace how attention targets constraint bottlenecks (Minimum Remaining Values), MLP layers pick optimal actions, and dead-end states cleanly trigger backtrack commands. In the second part, we pivot to macro-capability scaling with INFUSER, a self-evolution framework where a Generator and Solver co-evolve using unstructured document pools. Rather than relying on simple difficulty heuristics, INFUSER uses an optimizer-aware influence score and a novel RL objective to ensure the generator synthesizes questions tailored to the solver’s immediate learning needs. The resulting framework yields significant gains on complex math and coding benchmarks, demonstrating that an $8\text{B}$ co-evolving generator can outperform frozen $32\text{B}$ models.
Bio: Prof. Zhuoran Yang is an Assistant Professor of Statistics and Data Science and Computer Science at Yale University, affiliated with the Yale Institute for Foundations of Data Science and the Center for Algorithms, Data, and Market Design. His research focuses on machine learning, reinforcement learning, multi-agent systems, game theory, optimization, and the foundations of artificial intelligence, particularly the emergent behaviors of large language models. Previously, he was a postdoctoral researcher at UC Berkeley, working with Michael I. Jordan. He received his Ph.D. from Princeton University, advised by Jianqing Fan and Han Liu, and his bachelor's degree in Mathematics from Tsinghua University in 2015.