@LiorOnAI
Small models just beat giant LLM agents at their own job. Not by thinking harder, but by coordinating better. A new system just outscored GPT-5 on Humanityโs Last Exam, using far less compute. ๐ง๐ต๐ถ๐ ๐๐๐๐๐ฒ๐บ ๐ฟ๐ฒ๐ฝ๐น๐ฎ๐ฐ๐ฒ๐ ๐ผ๐ป๐ฒ ๐ฏ๐ถ๐ด ๐ฏ๐ฟ๐ฎ๐ถ๐ป ๐๐ถ๐๐ต ๐ฎ ๐ฐ๐ผ๐ป๐ฑ๐๐ฐ๐๐ผ๐ฟ Instead of one model doing everything, it assigns roles. โข Large models handle hard reasoning. โข Small models handle routine steps. โข A controller decides what to call, when. That controller is trained only to make decisions. ๐๐ ๐น๐ฒ๐ฎ๐ฟ๐ป๐ ๐ฐ๐ผ๐ผ๐ฟ๐ฑ๐ถ๐ป๐ฎ๐๐ถ๐ผ๐ป, ๐ป๐ผ๐ ๐ฝ๐ฟ๐ผ๐บ๐ฝ๐ ๐๐ฟ๐ถ๐ฐ๐ธ๐ It uses reinforcement learning, not hand rules. Rewards optimize three things at once: - Task success - Latency - Compute cost ๐ง๐ต๐ฒ ๐ฟ๐ฒ๐๐๐น๐๐ It scores 37.1% on HLE versus 35.1%. Runs about 2.5ร faster. Uses roughly 70% less cost. This lets you build agents that scale by coordination, not parameters.