@abhayesian
New Anthropic Fellows research: Alignment auditingāinvestigating AI models for unwanted behaviorsāis a key challenge for safely deploying frontier models. We're releasing AuditBench, a suite of 56 LLMs with implanted hidden behaviors to measure progress in alignment auditing. https://t.co/JNShb62b8y