AI

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Researchers at OpenAI tweaked two settings in their GPT-5.6 model to improve its performance on the ARC-AGI-3 benchmark. By enabling 'reasoning' and 'compaction', they were able to triple their scores. The exact nature of these settings is not specified, but it appears that they allow the model to process information more efficiently. This improvement was noted in both scoring and computational efficiency. OpenAI does not provide further details on how exactly these tweaks wo
Researchers at OpenAI tweaked two settings in their GPT-5.6 model to improve its performance on the ARC-AGI-3 benchmark. By enabling 'reasoning' and 'compaction', they were able to triple their scores. The exact nature of these settings is not specified, but it appears that they allow the model to process information more efficiently. This improvement was noted in both scoring and computational efficiency. OpenAI does not provide further details on how exactly these tweaks worked. --- Why it matters: This matters because it shows that even small adjustments can have significant impacts on AI performance. Researchers may be able to apply similar tweaks to improve their own models' scores on this benchmark, which is a key measure of general intelligence. Source: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores

This article was originally published at: https://openai.com/index/how-two-settings-tripled-our-arc...