Manifestation Units — standardizing neural-network component analysis
Mechanistic interpretability produces detailed component analyses, but outputs stay siloed: circuit diagrams, feature lists, selectivity tables locked in per-study notebooks.
New work proposes Manifestation Units, a typed tuple protocol (E, S, R, D, G) plus attention-head primitives (T) for transformers, to standardize how component statistics are represented and made queryable downstream.
Treats the representation layer itself as the bottleneck, decoupled from analysis methods, enabling reusable, composable knowledge for audits and interventions.