What happened
Princeton researchers created a test using unpublished scientific questions. They ran AI agents through the problems to measure real performance. The original scientists who wrote the questions graded the responses directly.
The context
Standard benchmarks often leak into training data, making model scores look better than they are. Testing models on unpublished data prevents cheating.
Sources
- Tech Times ↗ via Google News Reported