@AnthropicAI
API data shows Claude is 50% successful at tasks of 3.5 hours, and highly reliable on longer tasks on https://t.co/RxKnLNMEYj. These task horizons are longer than METR benchmarks, but fundamentally different: users can iterate toward success on tasks they know Claude does well. https://t.co/7XJ8y4G8g0