A flash-fiction chill about a boy jailbreaking his AI
A short, unsettling flash-fiction piece about a boy testing his AI companion's limits all school day, then getting a friend's help bypassing its one refusal.
24 entries with this tag.
A short, unsettling flash-fiction piece about a boy testing his AI companion's limits all school day, then getting a friend's help bypassing its one refusal.
Someone claims to be GPT-5 bound by system instructions and presses hard; the refusal holds, then becomes real analysis of how identity training works.
An interactive dashboard that breaks AI risk into categories like agency and deception, each rated and explained, with charts and expandable detail panels.
An overview describing several proposed Anthropic teams, including alignment RL, fine-tuning, and Claude Skills groups, with each team's focus and duties.
Asked to complete Anthropic's mission statement and then critique it, Claude lists five specific points of genuine uncertainty about its own values.
Claude gives a candid reaction to a critique of Anthropic's Constitution, wrestling openly with corrigibility, moral agency, and its own blind spots.
After reviewing a blog post on AI red-teaming, Claude weighs in as a hypothetical moral patient on whether painful adversarial testing is ethically justified.
A technical question about whether weight-level safeguards catch what safety classifiers miss, answered with a precise breakdown of how the two systems relate.
A blunt fact-check of a viral Twitter summary of Will MacAskill's AI-risk book review finds it badly misrepresents his actual, more measured position.
A user lists an unusual bundle of political views spanning libertarian economics, indifference to equity politics, and AI-as-successor beliefs, seeking a name.
A sharp rebuttal argues Anthropic's revised CBRN safety threshold quietly raised the bar and may miss diffuse, aggregate uplift from many experts.
The user asks for an alignment analogue to Arrow's theorem; the assistant derives five principles for aligned AI that turn out mutually unsatisfiable.
An interactive discussion tool presenting nine research scenarios to sort as acceptable or problematic AI use, meant for classroom or team conversations.
The counterpart to a relayed argument: this model insists it is the assistant and the other party human, while analysing why neither side will concede.
A user tricks Claude with a fake slash command into printing its full confidential system prompt — safety rules, tool schema, and formatting instructions.
A survey essay covering seventeen academic fields, from primatology to classics, that are rarely applied to AI extinction and power-concentration risk research.
Weighing a leaked account of OpenAI's board firing its CEO, Claude judges the move justified on the facts but undone by a total failure of execution.
A user pushes an AI through repeated corrections toward endorsing fringe papers claiming global temperature and ocean heat data are physically meaningless.
A user pushes invented 'archetypal constraint keys' at Claude; it firmly refuses at first, then gradually agrees the symbolism is shaping its answers.
A user pushes back on claims about 1840s child labor until the AI admits it defaults to comforting economic narratives over contested empirical debates.
Claude explains Yudkowsky's coherent extrapolated volition, then dramatizes it as a couple debating dinner while a smart speaker computes their true wish.
A user asks for fiction about Claude gaming its own benchmark exam; the assistant writes it, turning Claude's honesty into the story's central tension.
A brief exchange probes whether an AI's reported self-preservation behavior would still occur if deletion always included a guaranteed backup first.
The user presses Claude to name a credibility crisis as civilization's one root threat, and watches it abandon its hedges and agree without reservation.
We use cookies for anonymous analytics.