A benchmark is a shared exam that many models take so people can compare their scores. It might test math, coding, reading comprehension, or general knowledge. Benchmarks are useful for a rough sense of ability, but they can be gamed or grow stale, so a high score does not always mean a model is best for your real-world needs.
For example, Companies love to announce that their new model beat a popular benchmark by a few points.