@davis7
I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh... gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80% (was a near miss on the "x"s so actually over 80%) I am very confused https://t.co/NdDSTjoHLj