ACMS Statistics Seminar - Yue Xing, Michigan State University

-

Location: 154 Hurley Hall

Yue Xing

Michigan State University

3:30 PM
154 Hurley Hall

Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

Alignment is the final stage in the training pipeline of large languagemodels (LLMs), aimed at instilling capabilities beyond language comprehensionand instruction-following that are learned during pretraining and instructionfine-tuning. While various studies have developed and enhanced alignmentalgorithms for different tasks and data, there lacks a unified understanding onhow and why these algorithms work. In this work, we provide a unifiedanalysis framework from a divergence-based perspective and consider generalalignment settings, such as reinforcement learning with verifiable rewards(RLVR) and Preference Alignment (PA). Within this unified framework, weidentify some potential issues in existing alignment methods, such as thewidely used Group Relative Policy Optimization (GRPO). We further propose$f$-GRPO, a class of on-policy reinforcement learning, and $f$-Hybrid AlignmentLoss ($f$-HAL), a hybrid on/off policy objectives, for general LLM alignmentbased on variational representation of $f$-divergences.

View Poster