@chatgpt21
GPT-5.5 Scores .43% on ARC AGI 3! - GPT-5.5: 0.43% - Opus 4.7: 0.18% - GPT-5.4: 0.20% - Claude 4.6: 0.45% - Gemini 3.1: 0.4% The reported failures for GPT 5.5 were: - True local effect, false world model - Wrong level of abstraction from training data - Solved the level, didn’t reinforce the reward I think the full analysis will help OpenAI have a well rounded understanding of where the models are failing in certain modalities