Six alignment algorithms dissected layer-by-layer inside language models
arXiv paper opens the black box on post-training alignment: mechanistic analysis of PPO, DPO, SimPO, ORPO, GRPO, and KTO across three model families.
Using layer-wise probing, Sparse Autoencoders, and crosscoders, researchers localized where preference signals live and mapped how each method reshapes latent geometry differently.
KTO and GRPO stand out: they boost linear separability through constructive feature sharing, while others induce distinct representational shifts.