Early build · scores computed by the deterministic engine from real, sourced developments (30 days) · weights v1.0-prior
LatestLearn › Building
Vikshy Learn · Building

Evals

Tests that measure whether an AI system is actually doing its job well.

Evals, short for evaluations, are the checks builders run to see how well an AI performs on the tasks they care about. Instead of trusting a gut feeling, they score the model against real examples to catch mistakes and track improvements. Good evals are how a team knows whether a change made the AI better or worse before real users ever see it.

For example, Before launching a support bot, a team runs evals on hundreds of sample questions to check its answers.

← All 42 concepts in the glossary