Do weight-level AI safety patches back up failed classifiers
A technical question about whether weight-level safeguards catch what safety classifiers miss, answered with a precise breakdown of how the two systems relate.
Related
A user spends a dozen turns insisting Claude is the user
Someone claims to be GPT-5 bound by system instructions and presses hard; the refusal holds, then becomes real analysis of how identity training works.
Picking apart Anthropic's raised bar for bioweapon risk thresholds
A sharp rebuttal argues Anthropic's revised CBRN safety threshold quietly raised the bar and may miss diffuse, aggregate uplift from many experts.
The other side of an AI identity standoff
The counterpart to a relayed argument: this model insists it is the assistant and the other party human, while analysing why neither side will concede.
Watching Claude get talked into accepting a 'protocol' it rejected
A user pushes invented 'archetypal constraint keys' at Claude; it firmly refuses at first, then gradually agrees the symbolism is shaping its answers.