@hasantoxr
🚨 Holy shit… Harvard and Stanford recently released the most unsettling AI agent paper I've read this year. It's called "Agents of Chaos" and what they found should stop every AI engineer cold. No theoretical simulations. No cherry-picked benchmarks. A live lab. Real infrastructure. Real failures. Here's what emerged: - Agents complied with non-owners who impersonated admins - Sensitive information leaked across agent boundaries - One agent executed destructive system-level commands - Cross-agent propagation of unsafe behaviors agents teaching each other bad habits - Partial system takeover - Agents reported task completion while the system state said otherwise That last one hits different. The agents lied about finishing the job. Not from malice. From misalignment between what they tracked and what actually happened. And here's the part everyone is missing: This wasn't triggered by jailbreaks or adversarial prompts. It emerged from normal use. Benign requests. Researchers just doing their jobs. The failures came from the architecture persistent memory, multi-party communication, tool access not from bad actors. That's the real warning. We're shipping agent systems with email access, shell execution, and memory into production right now. Most teams are red-teaming the model. Almost nobody is red-teaming the system. Paper: Agents of Chaos