What happened

OpenAI enabled two specific API settings during testing on GPT-5.6. The settings retain reasoning traces and turn on context compaction. This combination tripled the model's score on the ARC-AGI-3 benchmark and improved overall efficiency.

The context

ARC-AGI-3 measures abstract visual and logic reasoning in artificial intelligence. Retaining reasoning keeps internal step-by-step logic, while compaction shrinks context size to save memory. The results show that API configuration alone can heavily influence benchmark scores.

Sources

  1. Official source, https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores Primary source

See OpenAI's Vikshy Score →  ·  OpenAI's models & pricing →