Emotion Geometry Across Scale
Multidimensional scaling (MDS) projections of 171 emotion fiction vectors across Gemma 2 scales. Top row: base models (2B, 9B, 27B); bottom row: instruction-tuned (IT) variants. Each point represents one emotion vector derived via the Sofroniew contrastive method (mean fiction stories − mean neutral stories, PCA-denoised, unit-normalized). MDS was computed on cosine distance matrices. Despite a 2× increase in hidden dimension (d=2304 to d=4608) and divergent training histories (2T–13T tokens; 2B and 9B distilled from 27B teacher), the relative geometry of the emotion manifold is conserved across all six models (RSA ρ ≥ 0.969 for all 15 pairwise comparisons). Positive-valence emotions (happy, delighted, excited) consistently cluster opposite negative-arousal states (tired, depressed, bored), with threat-related emotions (alarmed, afraid, angry) forming a distinct region. Comparing vertically within each column, RLHF preserves angular structure while amplifying vector magnitude (IT/base norm ratios: 1.10×, 1.07×, 1.31×), with the strongest amplification at the 27B scale. The geometry is a conserved feature of the architecture family.
LLMs often use contrastive prose in writing: “It’s not X. It’s Y”. Is this a stylistic tic? Or is it a behavioral artifact of post-training? Preliminary findings show this “negative parallelism” is absent in base models and SFT-only outputs. The pattern emerges predominantly in DPO trained models and is observed in DPO + RLVR trained models as well. We are further exploring whether contrastive post-training reshapes latent geometric structure and conceptual relationships in ways that affect intermediate representations, output distributions, and syntactic selection during generation.