Constructive Alignment — reframe AI safety as managing preference drift
arXiv paper challenges the standard alignment assumption that human preferences are fixed targets. Instead, preferences are dynamic and co-constructed through interaction with AI systems.
Proposes “Constructive Alignment” — treating alignment as a control problem over evolving preference trajectories, not static satisfaction. Draws on behavioral economics, psychology, and constructivist social theory.
Key insight: as AI becomes more personalized and socially embedded, it increasingly shapes what people attend to, value, and endorse over time—making preference trajectory management a first-order safety concern.