Alignment research must shift from frozen weights to constant weight updates under continual learning
Current alignment focuses on ensuring frozen weights behave well during deployment. With continual learning, weights update constantly, requiring new research on preventing jailbreaks, dece…