Raw model completions behind the v3.1 report. Amplifying or negating a personality-trait
LoRA (the weight control) shifts what a model says about itself and how it handles
unsafe requests. Filter or search the text, then draw random samples. Full corpus — all
176K samples. Text search scans every completion, so it takes a few seconds.
Judgments dataset: Butanium/persona-negation-judgments · filters are kept in the URL and restored on reload — reset