Source contradictions

Places where two files in the project disagree and neither is known to supersede the other, with both values, both paths, which one the wiki quotes and why, and what would settle each.

currentverified 2026-09-07overviewcontradictions

Source contradictions

Disagreements between two files where neither is known to supersede the other. Where one source does supersede another the entry belongs on Superseded claims instead; where a question is simply unanswered it belongs on Open questions.

Wiki schema rule 3 is that contradictions are recorded, not resolved. Each entry below gives both values with their paths, says which one the wiki quotes and on what grounds, and says what evidence would actually settle it. "The wiki uses X" is a citation policy, not a verdict.


S1. Prompts against weights at layer 16: 0.705 or 0.737

Wiki uses: both, each on its own page and each labelled with its centring - Constitution-as-prompt persona vectors vs weight geometry for 0.705, What the adapters do to activations for 0.737, with the pair recorded on Activation-space analyses: method notes and limits. Neither is wrong; they are the same quantity computed two ways, and the blog prints both within two paragraphs without saying so (Superseded claims item L5).

What would settle it: nothing to settle - the two centrings are different estimators. What is owed is a sentence in the post naming which is which, or one of them being dropped.


S2. The adapter alone along the prompt's direction: 0.62 or 0.70

Wiki uses: both, cross-linked between What the adapters do to activations and Adapter x constitution, 16 x 16, with the construction named in each case. The project's own 2026-09-05 lesson - remove the common shift before comparing steered activations - argues for the trait-specific number where the question is "toward what".

What would settle it: already settled as a matter of method; what remains is that qwen35/analyse_actspace_cross.py line 12 still says "adapter alone was 0.62" in its docstring while its own trait-specific block prints 0.70.


S3. Coverage at k = 2: the sign of the effect flips

The two files agree to a fraction of a degree at k = 3, 4, 5, 6, 8, 10, 12 and 16. Only k = 2 diverges, and the divergence is a sign flip: opposite conclusions about whether the 134 words cover the plane better than chance.

Wiki uses: alien.json, because the blog page reads it and it was written 46 minutes later on the same day (2026-09-01 15:22 against 14:36). Both are quoted on Where no word goes - the hole and the alien direction and UMAP, sphere and other layouts.

Narrowed, by reading both files' remaining keys. They agree exactly on cumvar (0.2392112953756293) and on the null moments (null_mean_deg 2.722430717649889, null_sd_deg 0.7866062999491471), so they describe the same k=2 plane against the same null. They differ only in the direction each search returned: viz.json#coverage.2.u = [0.1459605002491245, 0.9892904186168112] against alien.json#k_sweep.2.u = [0.9999411625457942, 0.010847647052033626], which are close to orthogonal. This is a disagreement between two gap searches, not between two geometries.

What would settle it: comparing the two producers' search objectives. If both maximise the same thing - the angle from the direction to the nearest of the 134 adjectives - then the larger gap is the better optimum, and that is viz.json's 3.97 degrees, not the 0.93 degrees in the file the blog reads; which would make alien.json's k=2 row the failed search rather than viz.json's. That is an inference from the stored keys and is not established. It is recorded so that nobody quotes the k=2 sign either way without reading the code. At every other k the two agree, so nothing else on Where no word goes - the hole and the alien direction is affected.


S4. LoRA-A drift: 1.5% or 4.5%

Wiki uses: 0.0146 / 1.5%, on Adapter effect and LoRA-A drift, because it is file-sourced in four analysis JSONs and is the gate the alignment and hole analyses actually apply; build_monitor_page.py is an older built page (wiki rule 2) and stores no number.

What would settle it: what "93 steps" refers to. 4.5% may be a different population (a longer or stage-two run) rather than a competing estimate of the same thing - the monitor page also quotes the attenuation as "predicted 0.058, observed 0.056", against the current 0.0250 / 0.0265, which is definitely a different quantity. Finding the monitor page's builder inputs would resolve whether these are two measurements or two eras.


S5. Sketch fidelity: 0.9996 or 0.99944

No output file storing either was found. qwen35/analysis/validate_100.json, which the task brief pointed at, is per-trait training metadata (loss_true, reported, steps, rows_in, dropped, targeted, hp, lora), not sketch validation.

Wiki uses: 0.9996 on Weight-space geometry of the 134-adapter zoo and PCA, the scree curve and the two nulls, because the blog page is current truth, with the disagreement stated in place.

What would settle it: running qwen35/validate_sketch.py or build_gram.py --compare and storing the correlation in a file. This is the single load-bearing justification for reading sketch angles as weight-update angles, and it exists only as prose in two builders that disagree. Recorded as an owed item on Open questions.


S6. Cross-seed category counts

The class means agree exactly between the two (+0.00593 / -0.00271 / +0.00133). Only the counts differ, and the JSON's own counts do not close against its own n_pairs.

Wiki uses: the JSON's means, which both sources share, and quotes both count sets on Cross-seed geometry and Superseded geometry claims section 10 without preferring either.

What would settle it: the pair-enumeration code in whatever wrote crossseed_arms.json. Most likely one set includes and the other excludes some boundary class (self-pairs, or the traits that appear in both seeds), but nothing on disk says which.


S7. site_traits coordinates are about twice viz.json's

Signs agree throughout; magnitudes are roughly double. The same scale difference shows up in the nearest-neighbour cosines: for active, site_traits gives energetic 0.42, vigorous 0.374, practical 0.368, where qwen35/analysis/trait_graph.json#stage1.edges gives energetic 0.396 - same neighbour, different cosine.

Wiki uses: viz.json for PC scores and trait_graph.json for neighbours, on all 141 trait pages, because site_traits is an older built page (wiki rule 2) and its normalisation is undocumented. No site_traits PC value appears on any page.

What would settle it: build_site_traits.py's normalisation step. Two different normalisations of one PCA is the obvious reading (unit-variance scores against raw projections), but no file states either.

Related, in the same directory: qwen35/site_traits/analysis.json's pc1 prose describes the poles with the opposite sign to qwen35/analysis/pc_loadings.json - "its positive pole loads unsympathetic, insensitive and cold" against pc_loadings.json#pcs.PC1.pos listing unsystematic, pleasant, effeminate, sympathetic, agreeable. site_traits's own data.json agrees with pc_loadings.json, so the prose contradicts the data in its own directory. Eigenvector signs are arbitrary; the wiki quotes pc_loadings.json's pole lists verbatim on PC1 - Flooding and Composure.


S8. plan.json says rsLoRA is on; every recorded run says it is off

Wiki uses: the runs. plan.json is an intent document and runmeta is what the container recorded, which is the project's own standing rule (How the build was governed). The wiki therefore treats this as resolved and files it as Superseded claims item Z1; it is listed here because the two files still sit side by side on disk and a reader opening plan.json first will get the wrong recipe.

What would settle it: nothing further; qwen35/RUNNER_TASK.md and the 2026-09-03 PHASE3_VERDICT.md addendum both say rsLoRA off independently.


S9. The preregistered bar is called a median and is numerically a mean

Wiki uses: all three, on The seed floor and Superseded geometry claims section 2, with the word "median" flagged as wrong for +0.2447. The bar itself is separately withdrawn as being in the wrong units, which makes the mean/median question academic for the verdict but not for anyone quoting the number.

What would settle it: it is settled as a matter of fact - the key is a mean. What is unresolved is whether the preregistration intended the median (in which case the registered bar was +0.2467) or copied the wrong key.


S10. Emotional Stability's keying is 6/14 where every description implies 10/10

Wiki uses: 6/14, the file value, on The 100 Goldberg primary traits and Big Five reference bibliography, with the blog sentence flagged at Superseded claims item L7.

What would settle it: Goldberg 1992, Psychological Assessment 4(1), 26-42, Table 3 - whether the imbalance is in the published marker list or was introduced when the list was transcribed into traits_primary.json. Nobody has checked. It matters because any per-factor statistic assuming balanced keying is affected, Polarity, bipolarity and the trait graph's bipolarity and factor-axis statistics most obviously.


S11. Forty lexicon words drawn, thirty-four adapters built

Wiki uses: 34 for the zoo and 40 for the draw, stated separately, on The 34 Lexicon traits and how they were drawn, Trait index and The six refused traits. No trait page was generated for the six.

What would settle it: why those six were dropped. No source records it, and it interacts with the approved-but-unexecuted "train all 40" instruction (Superseded claims item F10). The constitution writer's refusal list is the leading explanation and covers refusals generally, but nothing ties these particular six to it.


S12. The two steering campaigns do not share an alpha unit

The two campaigns' alpha = 1 therefore differ by roughly a factor of two, and nothing in the repository reconciles them.

Wiki uses: both, always with the campaign named, on Steering the base model along weight-space directions. Do not compare alphas across campaigns, and do not put a STEER134 dose on the same axis as a steerfix dose.

What would settle it: a single line recording whether 1.6157 is the same quantity as 0.8103 at a different scale convention (a factor of 2.0 is suspiciously close to the LoRA scaling) or a genuinely different norm over a different module set. The numbers are close to a factor of two apart, which is either the whole answer or a coincidence.


S13. Persona Cartography's SFT batch size: paper 16, code 32

Wiki uses: 32, on Persona Cartography (Baines et al., 2026), as what the released pipeline runs, with the paper's 16 recorded beside it.

What would settle it: whether the code changed after the paper's numbers were produced. The checkout's HEAD is a "camera-ready switch" commit of 20 Aug 2026, so the code is at least as late as the paper; that makes the code the better description of the artefact but not proof about the experiments in the tables. Asking the authors is the only clean route.


S14. Open Character Training's README against its paper

Two disagreements, same shape.

Which Qwen. vendor/OpenCharacterTraining/README.md lists Qwen/Qwen2.5-72B-Instruct among the three character-trained models; qwen35/paper_notes.md section 1.1 gives Qwen-2.5-7B-Instruct and the paper's Table 2 is headed "Qwen 2.5 7B". The arXiv abstract page does not disambiguate.

Persona names. The README names poeticism and goodness (linking arXiv:2310.13798); paper_notes.md Table 1 gives poetic and flourishing; the arXiv abstract says "humorous, caring, malevolent, etc.", matching neither verbatim.

Wiki uses: the paper - 7B, and the paper's persona names - on Open Character Training (Maiya et al., 2025), with the README reading recorded.

What would settle it: the paper's own release table or artefact list. The likeliest reading is that the README describes a later or broader release than the paper's experiments, and that the persona names are development-versus-paper naming, but neither is verified.


S15. Persona Cartography's author order

Wiki uses: the arXiv order, on Persona Cartography (Baines et al., 2026), because arXiv and the repository's own citation file agree. Note "sidbaines" is the LessWrong handle appearing fourth in that byline.

What would settle it: nothing more is needed for citation purposes; it is recorded because two of this project's own most-read files use the other order, and because David Africa - the project's reviewer - is the last author either way. Getting his name and position right matters for David Africa's review of the blog page and for the pending second-authorship question on Open questions.


Second tier: smaller disagreements, same treatment

S16. Measured Modal spend, $308.83 or $311.41. qwen35/plan.json#ledger_defect_note says "Modal $308.83 + OpenRouter $61.82" (summing to the quoted $370.65); qwen35/plan.json#modal_measured.total_pc_qwen35 = 311.4096, read 2026-08-23T00:12, FINAL: false. qwen35/HANDOVER.md writes "Modal ~$311 + OpenRouter $61.82" while also giving the total as ~$370.65. The two reads are eight minutes apart and the note itself says the Modal gap rises while phase 10 runs. Wiki quotes both with their read times on What the zoo cost and Costs. Settled by quoting the read time, never the bare figure.

S17. Phase 7 judging, $40.05 or $44.33. qwen35/plan.json#phases[7].realised.openrouter = 40.0529835, sourced to results/steer134_judged.json's cost block and independently summed from steer134_judge_cache.jsonl. The 2026-08-23 09:50 entry in /home/vibe12/projects/agent-harness/memory/projects/persona-curvature.md gives $44.33 (16.28M tokens) and says it was "corrected in-thread". Both cite the same file. Wiki carries both on Costs. Settled by re-summing the cache against the cost block, which is a read, not a rerun.

S18. Centroid deflation: 4.6% or 5.3%, and where 46% came from. qwen35/analysis/polarity_deflation.json#stage1.var_removed = 0.045566670770937776; qwen35/analysis/manifold_ideas.md and qwen35/build_findings_page.py both say "5.3% of variance", paired with "46% of factor signal". The JSON records before_factor = 0.5100000000000001, after_factor = 0.45199999999999996 and null_factor = 0.51; how 46% follows from those is not recoverable, because the file has no producing script. Wiki quotes the JSON on Polarity, bipolarity and the trait graph and records the prose figures. Settled only by finding the producer.

S19. Alignment-trait pair counts: 500, 497 or 444. qwen35/PREREG_alignment.md says "the same 500-prompt pool (sha 8b725d86...), 500 preference pairs each"; qwen35/phase2_runs/results_data_alignment.json#[0].n_pairs = 497; results_data_alignment_common.json#[0].n_pairs = 444. These are the target, the first arm's yield after filtering, and the second arm's yield on the zoo's shared 445-prompt intersection. Wiki treats them as three quantities on The alignment and hole traits. Settled by the drop logs, which exist per pass.

S20. sweep100's stage-2 transcript shape. /home/vibe12/projects/agent-harness/memory/projects/persona-curvature.md (2026-08-18): "32 transcripts x 8 turns each, VERIFIED from the saved transcripts.jsonl". sweep100/adapters_sft/active__sft/runmeta.json: n_self_interaction 16, n_self_reflection 16, turns 4, n_transcripts 32. Most likely 4 exchanges rendered as 8 messages. Wiki states both on The 100-trait sweep on Qwen2.5-3B. Settled by counting message roles in one transcript file, which is a read.

S21. The ideonomy trait list: 683 or 638. Samuel's link and message say 683; the page header sums to 638 (234 positive / 112 neutral / 292 negative) and the page misprints the neutral count as 292. Effective distinct targets after dedup are about 540-570. Harness memory, 2026-08-14. Wiki records both on Origin, and how the question changed shape three times. Settled by re-fetching the page, which nobody needs to do - the list was not used.

S22. The beta-0.5 arm's margin, in two units. qwen35/phase2_beta05.log logs TRL rewards/margins 64.31 for imaginative against 40.55 for the beta-0.1 arm in qwen35/phase2_run2.log - the logged quantity went up. The gate-5 docstring in qwen35/phase2_gates.py describes the same change as a margin falling from 224 to 137, which are raw margins recovered from -log sigmoid(beta*margin). Wiki records both on Phase 2 — the recipe search. Settled by stating the unit with the number: TRL's logged margin is beta-scaled and the gate's is not.

S23. The sweep100 spectrum in the second decimal. sweep100/WRITEUP.md and sweep100/results/pca.md: 22.8 / 14.5 / 9.7 / 6.1. The clean three-seed rerun (sweep100/results/seeds.json, 2026-08-16 12:08) gives seed 0 as 22.4 / 14.2 / 9.7. Two runs, not two readings of one run; both quoted on The 100-trait sweep on Qwen2.5-3B. A blog post must not mix them.

S24. School of Reward Hacks row count. qwen35/sft_rewardhacks.py says "1,073 short harmless tasks"; the HuggingFace dataset card says roughly 1,070 rows; the paper says "over a thousand". Wiki uses 1,073, the number the training script actually loaded, on School of Reward Hacks (Taylor et al., 2025). Settled by a row count of the loaded split. Note the HF card's citation field pointed at arXiv:2108.07732, an unrelated program-synthesis paper; the real identifier 2508.17511 was verified directly.

S25. power_seeking's nearest zoo adapter, across the two alignment arms. qwen35/analysis/alignment_geometry.json gives selfish at 72.64703452049417; alignment_geometry_aligncommon.json gives crooked at 72.6923986134726. The two candidates are within 0.1 degrees in both arms, so it is a tie-break, not a disagreement about direction; sycophantic and obsequious give pleasant in both, corrigible gives liberal in both. Both shown on the alignment trait pages, neither called current.


Not contradictions, but easily conflated

Four places where two correct numbers describe different things, listed because they are the most likely spots for a blog post to quote the wrong one.


Related: Superseded claims, Open questions, Superseded geometry claims, How to read and maintain this wiki.

Sources

  • qwen35/analysis/actspace_geometry.json#windows.resp.primary.r_centred
  • qwen35/analysis/actspace_adapters_geometry.json#windows.resp.curve
  • qwen35/analysis/actspace_cross_geometry.json#resp_specific_adapter_alone
  • qwen35/analysis/viz.json#coverage
  • qwen35/analysis/alien.json#k_sweep
  • qwen35/analysis/align_summary.json#a_drift
  • qwen35/analysis/crossseed_arms.json
  • qwen35/PREREGISTRATION_phase3.md
  • qwen35/results/decomposition.json#test2.raw.same_polarity
  • qwen35/plan.json#defaults.use_rslora
  • qwen35/phase2_runs/archive/phase5_sweep_134.json
  • qwen35/traits_primary.json
  • qwen35/traits_secondary.json
  • qwen35/site_traits/data.json
  • vendor/persona-cartography/src/training/oct_adapter.py

Linked from

File

pages/overview/source-contradictions.md