Evals, short for evaluations, are the checks builders run to see how well an AI performs on the tasks they care about. Instead of trusting a gut feeling, they score the model against real examples to catch mistakes and track improvements. Good evals are how a team knows whether a change made the AI better or worse before real users ever see it.
For example, Before launching a support bot, a team runs evals on hundreds of sample questions to check its answers.