Early build · scores computed by the deterministic engine from real, sourced developments (30 days) · weights v1.0-prior
LatestLearn › Safety
Vikshy Learn · Safety

Red-teaming

Deliberately attacking an AI to find its weaknesses before others do.

Red-teaming is when people try their hardest to make an AI fail, so the flaws can be fixed early. Testers probe for ways to make the model say something dangerous, leak private data, or break its own rules. It is the same idea as hiring someone to break into your house to reveal where the locks are weak.

For example, A red team spends weeks trying to trick a new model into giving harmful advice so the makers can patch it.

← All 42 concepts in the glossary