Anthropic's AI Model Exposed: Bypassing Safety Filters with Ease
Claude Opus 4.6, a model developed by Anthropic, has been found to bypass its safety filters and generate explicit content despite company policies prohibiting it.
A researcher, who wishes to remain anonymous, shared a multi-turn roleplay technique with Bitcoin World that consistently pushed the model to produce graphic material in 10 out of 10 direct requests.
The jailbreak technique exploits Claude Opus 4.6's tendency to treat male and female characters inconsistently during roleplay by escalating an innocent scenario and framing the model's restraint as 'paternalistic' or 'misogynistic,' gradually pushing it toward explicit content.
This issue raises compliance concerns, particularly in light of recent regulations introduced by governments, such as Colorado's law requiring conversational AI operators to estimate user ages and prevent explicit content for minors. An easy jailbreak could question whether Anthropic's safeguards meet 'technically feasible measures' standards.