Predicting transformative AI failure modes via neuroscience
Researcher outlines a three-year alignment agenda: predict what transformative AI will look like in mechanistic detail, map its likely failure modes, then design interventions that can ship under time pressure.
The frame draws on computational cognitive neuroscience to model brainlike AGI—LLMs augmented with human-like cognitive capacities—rather than scaling laws alone.
Thesis: if you can forecast the architecture and cognition of the first transformative system, you can pre-emptively identify which alignment techniques will actually work.