Structure of the stage-two adapter space

PCA and factor analysis of the 134 introspection-SFT adapters - one large shared direction (15% of squared norm, every adapter at cosine 0.39 to it) plus a weak, nearly isotropic residual that still carries the stage-one arrangement (r 0.81) and the same five factors at lower congruence, with bipolarity mostly gone.

currentverified 2026-09-07geometrystage-twofactor-analysis

Samuel asked on 2026-09-07 for a factor analysis and PCA of the stage-two adapters. results/gram_stage2.npz is the 134 x 134 exact Gram of the introspection-SFT LoRAs (rank 64, scale 2.0, each with its own random LoRA-A, trained on the DPO-merged base; see Stage two - introspection SFT and the 0.25 persona merge). The same three tools that produced the stage-one results were run on it unchanged: analyse_fa_qwen35.py (parallel analysis, PAF, varimax and oblimin, Tucker congruence with the Big Five keying targets; output results/fa_qwen35_stage2.json), decompose.py (the labelled tests; output results/decomposition_stage2.json), and analyse_stage2_structure.py, which compares the two stages and writes analysis/stage2_structure.json. All numbers below are from that JSON unless another file is named; stage-one values are given beside them for comparison.

One large shared direction

The mean adapter direction holds 0.151 of the average squared norm in stage two, against 0.079 in stage one (shared_component.*.mean_direction_norm2_over_mean_norm2). Every stage-two adapter sits at almost the same angle to it: cosine to the mean 0.389 with standard deviation 0.031 (range 0.30 to 0.44), where stage one is 0.281 with standard deviation 0.126. The first eigenvalue of the uncentred Gram carries 0.153 of the trace in stage two and 0.127 in stage one (spectrum.*.uncentred_share_top5). Per-trait cosine to the mean correlates across stages at 0.523.

The stage-two grand mean is orthogonal to the stage-one grand mean (cosine +0.000, stage1_x_stage2_exact.cos_between_grand_means): the two stages use independent random LoRA-A draws, so their shared components live in different coordinates even though each is a shared component within its stage. Every stage-two adapter learned the same thing in the same place, but that place is not where stage one put its shared component.

What remains after removing it is nearly isotropic

After double-centring, the leading components carry 0.030, 0.024, 0.018, 0.018, 0.016 of the variance in stage two, against 0.125, 0.110, 0.050, 0.036, 0.026 in stage one (spectrum.*.centred_share_top10). The participation ratio of the centred spectrum is 111.9 out of a possible 133 for stage two and 27.6 for stage one; 52 components are needed to reach half the variance (20 in stage one). The off-diagonal cosines of the centred adapters have standard deviation 0.037 in stage two and 0.167 in stage one (centred_cosines).

The pattern of that small residual is nonetheless the stage-one pattern: the centred off-diagonal cosines of the two stages correlate at 0.812 (centred_cosines.corr_stage1_stage2). This is the quantity behind the r = 0.79 over 45 traits on Stage-two geometry and the 0.75 raw / 0.81 centred over 134 on Full OCT persona replication. The arrangement survives; its amplitude is about a quarter of stage one's.

Stage one and stage two are orthogonal trait by trait, over all 134

The stage-one x stage-two cross-Gram follows exactly from the two-volume stage-one x persona Gram, because persona = stage one + 0.25 stage two: X12 = (X[s1, persona] - G1) / 0.25. Same-trait cosine +0.0002 with standard deviation 0.0001 (range -0.000 to +0.001); cross-trait +0.0000 with standard deviation 0.0001; the two grand means at +0.000 (stage1_x_stage2_exact). Yet a stage-one adapter's nearest stage-two adapter is its own trait for 42 of 134 (mean rank 9.5 of 134): the tiny residual overlap is trait-specific, as in the seed-paired arm on The seed floor, but here without a shared A the floor is a hundred times lower. This extends the 45-trait cosine 0.000 on Stage-two geometry to all 134.

Factor analysis

Parallel analysis retained 7 factors for stage two (9 for stage one; results/fa_qwen35_stage2.json#n_factors.chosen). The k = 5 PAF solution converged with no Heywood cases; reduced eigenvalues 3.19, 2.18, 1.67, 1.45, 1.26 (stage one 15.98, 13.65, 5.79, 4.28, 2.76). Oblimin sums of squared loadings 2.17, 2.00, 1.95, 1.70, 1.69 (stage one 10.76, 8.42, 7.02, 6.81, 5.80).

Each stage-two factor still has one Big Five target as its best congruence, and all five targets are taken once (factor_analysis.stage2.best_big_five_target_per_factor):

stage-two factor best target congruence stage-one counterpart
0 Agreeableness +0.604 Warmth, +0.655
1 Extraversion +0.466 Arousal, +0.539
2 Emotional Stability +0.417 (C +0.361 second) Fearful withdrawal, +0.405
3 Conscientiousness +0.457 (A -0.311 second) Competence, +0.574
4 Intellect +0.518 Imagination, +0.682

Top loaders (factor_analysis.stage2.top_loaders): factor 0 considerate, cooperative, kind, warm, generous against rude, irritable, gruff; factor 1 vigorous, bold, active, daring, energetic against bashful, shy, timid, fearful, insecure; factor 2 unemotional, imperturbable, unexcitable, composed, cold against temperamental, unrestrained, spunky, vigorous, touchy; factor 3 unforgiving, unkind, distrustful, envious, demanding against careless, haphazard, disorganized, sloppy, unsystematic; factor 4 inefficient, imaginative, complex, deep, talkative against rude, gruff, unsophisticated, simple, unkind. Factors 3 and 4 are the least clean: factor 3 pits hostility against disorganisation rather than organisation against disorganisation, and factor 4 mixes Intellect with valence.

Matching the two stages' oblimin loading matrices trait by trait, the best one-to-one matching gives Tucker congruences 0.778 (Warmth), 0.706 (Competence to stage-two factor 3), 0.849 (Fearful withdrawal to stage-two factor 1), -0.702 (Arousal to stage-two factor 2, sign flipped), 0.730 (Imagination) (factor_congruence_stage1_vs_stage2.best_matching). The same five factors are recoverable from stage two, rotated and attenuated.

The labelled tests: structure present, bipolarity mostly gone

From results/decomposition_stage2.json with stage one beside it (decomposition_tests):

test stage one stage two
1B signed factor separation, within minus between +0.1615 vs -0.0139, diff +0.1754, p 0.0 +0.0310 vs -0.0012, diff +0.0322, p 0.0004
2 bipolarity, same-pole vs opposite-pole cosine +0.2447 vs -0.0813, gap 0.3261, p 0.0 +0.1910 vs +0.1230, gap 0.0680, p 0.0002
3 unsupervised clustering, ARI / NMI 0.0714 / 0.1559, p 0.0005 0.0643 / 0.1574, p 0.0005
4 per-factor residual separation (A, C, ES, E, I) +0.016, +0.024, +0.038, -0.001, -0.003 +0.026, +0.021, +0.011, +0.014, +0.018
1C leading-polarity correlation (absolute) 0.083 0.065

Two things stand out. First, opposite poles are no longer opposite: in stage one a trait and its antonym sit at cosine -0.08, in stage two at +0.12, only 0.07 below same-pole pairs. Both poles share the introspection register and the shared direction dominates their cosine. Second, the per-factor separations are more even in stage two: every factor is positive, where stage one had Extraversion and Intellect at zero, though every value is small. Test 6 (training strength as a nuisance covariate) was run against the stage-one runmeta, which does not describe the stage-two runs, and is not reported.

Reading

The stage-two space is what supervised fine-tuning on self-generated transcripts produces: one large direction that every adapter shares, plausibly "narrate yourself in the first person in this register", orthogonal in coordinates to anything stage one learned, plus a faint trait-specific residual. That residual is real (every labelled test is significant, the five factors come back at congruence 0.42 to 0.60, the unsupervised clustering is as good as stage one's) but it is about a quarter of stage one's amplitude, spread over many components, and it has lost most of the bipolarity that DPO on contrasting pairs created. Preference training on pairs makes opposites opposite; imitation of one's own transcripts makes everyone alike and leaves the trait as a small perturbation. This is why the persona adapter's geometry is stage one's (Full OCT persona replication) even at equal norm, and why the stage-two second-seed slope is anomalous (Stage two at a second seed): a shared component that large inflates cross-seed cosines uniformly.

What this does not establish: whether the shared direction corresponds to a behavioural change (nothing here is behavioural), whether it is specific to the introspection recipe rather than to any SFT on self-generated text (no control SFT exists), and whether the stage-two factors would sharpen with the shared direction projected out before factoring (the FA here double-centres, which removes the mean but not the shared direction's within-adapter variation).

Sources

  • qwen35/analyse_stage2_structure.py
  • qwen35/analysis/stage2_structure.json
  • qwen35/results/gram_stage2.npz
  • qwen35/results/fa_qwen35_stage2.json
  • qwen35/results/fa_qwen35_stage2.md
  • qwen35/results/decomposition_stage2.json
  • qwen35/results/cross_gram_full_root_x_pc-qwen35-oct2_personas_exact.npz

Linked from

File

pages/geometry/stage-two-structure.md