EVENT DETAILS
Please join us in the Electrical and Computer Engineering Department at the Technological Institute for an hour-long seminar with Professor Lei Ying of the University of Michigan.
Abstract: Reward inference (learning a reward model from human preferences) is a critical intermediate step in Preference-based Reinforcement Learning (PbRL), such as Reinforcement Learning from Human Feedback (RLHF) for fine-tuning Large Language Models (LLMs). In practice, reward inference faces fundamental challenges such as distribution shift, reward model overfitting, and problem misspecification. An alternative approach is direct policy optimization without reward inference, such as Direct Preference Optimization (DPO), which offers a much simpler pipeline but only works in the bandit setting or for deterministic MDPs. This talk introduces new algorithms for stochastic MDPs and general preference models (link functions). The key idea is a sign-based policy perturbation approach based on zeroth-order optimization. We will discuss its applications to unknown link functions and federated RLHF.
Bio: Lei Ying is a Professor in the Electrical Engineering and Computer Science Department at the University of Michigan, Ann Arbor. He is an IEEE Fellow and an Editor-at-Large for the IEEE/ACM Transactions on Networking. His research focuses on the interplay between complex stochastic systems and big data, including reinforcement learning, large-scale communication/computing systems for big-data processing, private data marketplaces, and large-scale graph mining.
Location: Tech L440
TIME Thursday July 30, 2026 at 11:00 AM - 12:00 PM
LOCATION L440, Technological Institute map it
ADD TO CALENDAR&group= echo $value['group_name']; ?>&location= echo htmlentities($value['location']); ?>&pipurl= echo $value['ppurl']; ?>" class="button_outlook_export">
CONTACT Lee Onysko lee.onysko@northwestern.edu
CALENDAR Department of Electrical and Computer Engineering (ECE)