@EinsiaAI
The ultimate test for coding agents isn't local editing— it's whole-repo evolution, and right now, the survival rate is 5.4%. Today we’re releasing SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migration. Coding agents are getting very good at fixing bugs. But can they refactor an entire system, C → Rust, Maven → Gradle, POSIX → WebAssembly? We built 20 real migrations across projects, including SQLite, zlib, libsodium, and GraphHopper. 520 runs. Only 28 survived all 3 stages. 13/20 tasks were solved by nobody. System-scale migration is still wide open. Full breakdown 👇 GitHub: [https://t.co/xXyLQ3qq0C] Paper Link: [https://t.co/lO1Enh63q7] Einsia Website: [https://t.co/AGn0hF5gwL]