Red-teaming is when people try their hardest to make an AI fail, so the flaws can be fixed early. Testers probe for ways to make the model say something dangerous, leak private data, or break its own rules. It is the same idea as hiring someone to break into your house to reveal where the locks are weak.
For example, A red team spends weeks trying to trick a new model into giving harmful advice so the makers can patch it.