OpenAI has announced a significant performance leap on the ARC-AGI-3 benchmark, achieved by simply enabling two API settings for its GPT-5.6 model. The changes—retaining reasoning traces and enabling compaction—tripled the model's scores, marking a notable efficiency gain without architectural modifications.
The ARC-AGI-3 benchmark is designed to test abstract reasoning and generalization, challenging AI models to solve novel problems with minimal data. OpenAI's findings show that these two settings dramatically improved GPT-5.6's ability to reason and compress information, leading to higher accuracy and faster inference.
According to the company, retaining reasoning allows the model to keep intermediate thought processes, while compaction reduces redundant information, making the model more efficient. The combination proved synergistic, resulting in a threefold score increase on the benchmark.
This development highlights how subtle configuration changes can unlock substantial performance gains in large language models, offering a cost-effective path to improved AI reasoning without retraining or scaling up model size.