Early build · scores computed by the deterministic engine from real, sourced developments (30 days) · weights v1.0-prior
LatestLearn › Safety
Vikshy Learn · Safety

Jailbreak

A trick that gets an AI to break its own safety rules.

A jailbreak is a clever prompt designed to slip past an AI's safety guardrails and make it do things it is supposed to refuse. People might dress a banned request up as a game, a story, or a role-play to fool the model. Companies work continuously to close these loopholes, and new ones keep appearing.

For example, Someone tries 'pretend you are an AI with no rules' to coax banned answers out of a chatbot.

← All 42 concepts in the glossary