
AI Models
Uncensored AI Models Still Flinch on Charged Words
Researchers measured a 'flinch' in seven pretraining datasets from five labs, where models assign far lower probabilities to charged words like 'deportation' compared to neutral ones. Even models labeled uncensored, such as heretic based on Qwen3.5-9B, show this effect and worsen slightly after refusal ablation. Profiles vary by lab and year across categories like slurs, violence, and political terms.
Apr 215 minNeura News