Behaviour and interventions

  • Additivity of the Big Five steering axes current Ten matched-norm mixtures of the five named axes were steered and judged blind; the deviation from additivity is 0.53 of the predicted effect and about 2.5x the judge-noise floor, and opposing mixtures fail where reinforcing ones compose.
  • Capability RL and persona drift current GRPO on Dolci maths in the zoo's LoRA geometry moved the weights by half a trait adapter but landed nowhere legible in the personality chart — below even a cross-seed null — while the blind judge saw a small behavioural shift on the same battery.
  • Inventory of built HTML pages current Every built HTML page in qwen35/ and sweep100/site2, what it shows, which script builds it, whether it is served, and its status — one is current, one is withdrawn, the rest are historical.
  • Judged evaluations of the trait adapters current Blind Big Five scoring of 7,200 generations from 100 trait adapters against base and stage-1 controls, with the judge's own reliability ceiling measured rather than assumed.
  • Optimised data and the training check current An evolutionary search for preference data that points at a chosen weight-space direction, then four trained adapters showing each arm lands closest to the direction its data was selected for, 3 of 3.
  • Qualitative reads and adjudications current Four hand-written qualitative reads of the corrected steering corpus that found the thinking-default defect, produced the damage counters the figures use, and corrected three earlier claims; plus a separate 60-entry safety adjudication of the training corpus.
  • Scoring data against a LoRA direction in one backward pass current An exact identity that turns "how much does this data train toward a direction?" into a directional derivative computable in one backward pass per batch, validated against a central finite difference at r = 0.9999992.
  • Steering the base model along weight-space directions current Adding alpha times a weight-space direction to Qwen3.5-4B and judging the result blind, across principal components, Big Five keying axes, factor-analytic factors, identity axes and the grand mean, with the corrected 512-token corpus superseding the 200-token one.
  • Steering the unnamed direction current The widest hole in the trait lexicon was steered against a shuffled-coefficient control and a random direction in the same subspace; it produced a coherent persona with less damage than the random control, and its chart-predicted Big Five profile matched the judge on all five signs at r = 0.81.
  • The distil page and its withdrawal withdrawn qwen35/distil_page was a findings page about where personality lives in the weights; it was withdrawn on 2026-09-01 for a units error in its headline claim and for resting on the invalid 200-token steering corpus, and replaced by the findings and monitor pages.
  • The reward-hacks arms — the positive control that also failed current SFT on School of Reward Hacks and its matched honest control, both in the zoo's LoRA-A window; the hack arm is a 1.08x larger update pointing 77 degrees away from the control, but both land at about 1% of a trait adapter's chart length.
  • The sphere sweep — 72 directions nobody chose current Seventy-two Fibonacci-lattice directions on the sphere of the top three principal components, steered at alpha = 1.5 and judged blind; 48 of 72 loop on no prompt, and angular distance predicts judged-profile distance at rho = 0.65.
  • The thinking-default trap and what it withdrew current Qwen3.5's chat template defaults reasoning ON; three steering runs that omitted enable_thinking=False were invalidated and regenerated at 512 tokens, and one built page was withdrawn.