@MiniMax_AI
π Introducing OctoCodingBench, a new benchmark for aligned coding agents: https://t.co/oKaF7jjagb Passing tests β aligned behavior. An agent can produce code that aces every unit test while ignoring system guidelines, violating project conventions, or misusing tools. In real-world coding, how you solve matters as much as what you solve. Nobody wants an agent that ships perfect code while deleting your README, reformatting every file, and mass-commenting in LLM-ese. Don't let your coding agent paperclip-max your repo!