@nathanhabib1011
π¦ Claw-Eval π¦ π₯ @XiaomiMiMo's MiMo-V2.5-Pro at 1T π₯ @Zai_org GLM5.1 at 754B π₯ @XiaomiMiMo MiMo-V2.5 at 310B Congrats to @XiaomiMiMo for having 2 models in the top 3! The most impressive result though is @deepseek_ai with DeepSeek v4 flash a 210B model on par with models 4 times its size.. Super interesting bench from @_TobiasLee and team. Gathering tasks from @openclaw, PinchBench, OfficeQA, OneMillion-Bench, Finance Agent, and Terminal-Bench 2.0. This might be one of the more interesting and useful benchmark right now. Whether you are using @NousResearch's Hermes agent or @steipete's OpenClaw, you should choose your model according to real world tasks. https://t.co/8cq8gmgFz6