Activation space

  • Activation-space analyses: method notes and limits current Everything needed to reproduce the three activation-space runs of 2026-09-05, and the list of things they do not establish.
  • Activation-space analyses: overview current Three runs on 2026-09-05 asked whether a constitution used as a system prompt and an adapter trained on that constitution move Qwen3.5-4B's residual stream the same way, and whether the 134 shifts are arranged the way the 134 weight updates are.
  • Adapter x constitution, 16 x 16 current Sixteen adapters run under each of sixteen constitutions saturate rather than stack when they agree (1.18 against 1.00 for the prompt alone and 0.70 for the adapter alone) and split roughly evenly when they conflict (0.50 / 0.52); a first pass on uncentred vectors wrongly reported that the prompt wins 16 of 22 conflicts.
  • Constitution-as-prompt persona vectors vs weight geometry current The 134 constitutions used as system prompts produce activation shifts whose 134x134 geometry correlates with the adapters' weight geometry at r 0.705 at layer 16 (max 0.775 at layer 19), but the weight-space hole does not transfer.
  • Prompting versus training current Synthesis of the three activation-space experiments for the blog: prompting a constitution and training on it move Qwen3.5-4B's residual stream about equally far, along largely the same directions, into an arrangement that mostly matches the weight geometry, and when both are applied at once they saturate rather than stack.
  • What the adapters do to activations current Each stage-1 adapter, run with no system prompt, moves the residual stream as far as the trait's constitution does as a prompt (median ratio 1.01) and 62% of the way along the prompt's own direction, but the adapters' activation geometry tracks their weight geometry (0.865) more closely than it tracks the prompts (0.781).