@AnthropicAI
The checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled βInitβ below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents. https://t.co/YybDZfhA2Y