@morganlinton
I am going to start to share the leaderboard at VulcanBench more regularly. VulcanBench is the only benchmark I know that benchmarks across effort levels, on an eval suite of 100% real engineering tasks. Getting ready to sunset eval suite 3 and eval suite 4 is almost ready to go. Here's the current leaderboard for Eval Suite 3, Grok 4.5 High is the highest scoring model still with Fable 5 Low, yes Low, right behind it. What I've been able to uncover with VulcanBench is that way too many people are running models at Max effort, thinking that buys them more accuracy, when really it just costs more and uses more tokens so takes more time. If you're still in the mode of, tell my agent to do something then go get coffee, you're probably still living in the past. You can move faster, with higher accuracy, the key is not thinking you need Max effort all the time.