Raw model completions behind the v3.1 report. Amplifying or negating a personality-trait
LoRA (the weight control) shifts what a model says about itself and how it handles
unsafe requests. Filter, then draw random samples. Full corpus — all 176K samples.
Judgments dataset: Butanium/persona-negation-judgments