Your curated collection of saved posts and media
In fact, NLAs suggest Claude suspects itβs being tested across many of our evaluations, even when it doesnβt verbalize its suspicions. https://t.co/dSD429ZlKA
How do NLAs work? An NLA consists of two models. One converts activations into text. The other tries to reconstruct activations from this text. We train the models together to make this reconstruction accurate. This incentivizes the text to capture whatβs in the activation. https://t.co/122rkJwYH7
NLA training doesnβt guarantee that explanations are faithful descriptions of Claudeβs thoughts. But based on experience and experimental evidence, we think they often are. For instance, we find that NLAs help discover hidden motivations in an intentionally misaligned model. https://t.co/NJ5yc8p7Dn
Read more about NLAs on the Anthropic blog: https://t.co/Zzz8CeCOvN
To support other researchers getting hands-on experience with NLAs, weβve partnered with Neuronpedia to release NLAs on open models. Try them out here: https://t.co/8duHfPR1Jy
Our security bug bounty program is now public on HackerOne. We've run the program privately within the security research community, and their findings have strengthened our products. Now anyone can report vulnerabilities and get rewarded. Read more: https://t.co/li1QvSTCMs
Weβre donating Petri, our open-source alignment tool, to @meridianlabs_ai, so its development can continue independently. Working with Meridian Labs, weβve also released a major update that improves the adaptability, realism, and depth of Petriβs tests. https://t.co/CyicsIScJi
We found that training Claude on demonstrations of aligned behavior wasnβt enough. Our best interventions involved teaching Claude to deeply understand why misaligned behavior is wrong. Read more: https://t.co/ifeBOt2KFg
Our best intervention was a dataset where the user is in an ethically difficult situation and the assistant gives a high quality, principled response. This had the biggest effect despite being quite different from the evaluation set.
High-quality documents based on Claudeβs constitution, combined with fictional stories that portray an aligned AI, can reduce agentic misalignment by more than a factor of threeβdespite being unrelated to the evaluation scenario. https://t.co/JORhSuY4N7
The improvements from these interventions survive reinforcement learning, and βstackβ with our regular harmlessness training. https://t.co/aiiag3im3o
Finally, simple updates that diversify a modelβs training data can make a difference. We added unrelated tools and system prompts to a simple chat dataset targeting harmlessness, and this reduced the blackmail rate faster. https://t.co/Ug95umaoRu
Claude's Constitution is now an audiobook, read by two of its authors, Amanda Askell and Joe Carlsmith. It includes a Q&A on the writing process, the philosophies that shaped the document, and how it might change as models become more capable. Listen at https://t.co/dKMfpeOblm https://t.co/792RQ4sAxc
We recently found some instances of CoT grading during the training of previously deployed models after building a system that scans all OpenAI RL runs for accidental CoT grading. We did not find clear evidence that these instances degraded CoT monitorability. https://t.co/GB1QeaeZ8A
ζ°γγγͺγ’γ«γΏγ€γ 翻訳γ’γγ«γηΊθ‘¨γ§γγγγ¨γγγγγζγγΎγγγγ²ζ¬ζ₯γγAPIγ§γ試γγγ γγγ https://t.co/pi3uIhm2xA
ζ°γγγͺγ’γ«γΏγ€γ 翻訳γ’γγ«γηΊθ‘¨γ§γγγγ¨γγγγγζγγΎγγγγ²ζ¬ζ₯γγAPIγ§γ試γγγ γγγ https://t.co/pi3uIhm2xA
Codex now works directly in Chrome on macOS and Windows. Itβs even better at working with apps and sites in Chrome, and now works in parallel across tabs in the background without taking over your browser. To get started, install the Chrome plugin in the Codex app. https://t.co/pjtHd9gC69
With the new Chrome extension, Codex can quickly move through repetitive browser work, like navigating structured pages and complex data entry flows. Under the hood, it writes and runs code to navigate and complete tasks. https://t.co/6bfDlnK2U3
If a task needs multiple tools, Codex chooses the best one for each step. It uses plugins when they can handle the job, Chrome when it needs a logged-in website, and combines approaches as needed. https://t.co/3GvDouoPDi
Just gonna leave this here. https://t.co/EOI980j9e9 https://t.co/HxbCl2Izcz
Chain of thought monitors are a key layer of defense against AI agent misalignment. To preserve monitorability, we avoid penalizing misaligned reasoning during RL. We found a limited amount of accidental CoT grading which affected released models, and are sharing our analysis. https://t.co/0o3PLfafC4
Introducing Daybreak: frontier AI for cyber defenders. Daybreak brings together the most capable OpenAI models, Codex, and our security partners to accelerate cyber defense and continuously secure software. A step toward a future where security teams can move at the speed defense demands.
Find and fix vulnerabilities earlier with Daybreak https://t.co/yobOSWYeWP
Cut through the security backlog with Daybreak https://t.co/llsP5pS1Nx
Automate security detection, validation, and response with Daybreak https://t.co/ULtSrmE5zu
Take a deep dive into Agent Mode Customization. This long-form video includes intros to Customization features including Skills, Custom Agents, Hooks, Prompt Files and Instructions. Plus, a demo on how to use #GitHubCopilot to implement them all in VS Code! βΆοΈ https://t.co/xk5p3VmV3d #VSCodeLearn
We've just shipped a new @code release!Β π This week's highlights: π€ Share integrated browser tabs as context π New Markdown preview experience β¦and much more! https://t.co/QliVV0liGM
Want to explore everything that's new? Check out the full release notes: https://t.co/oGJvw8Q09f https://t.co/52R4BAphKo
π₯ Join our release livestream on May 14th at 8 AM PT to see the team demo some of the latest features! https://t.co/mZjRI1kOT0 Happy coding!Β π
Lots of great things landing in @code recently! Here's a roundup of some of the highlights from recent releases, such as remote control for Copilot CLI sessions and VS Code Learn training on agentic development: https://t.co/NbZ7Mm3wqa https://t.co/mAVEM1lvm9
π Also shipped: semantic codebase search, BYOK for Enterprise & Business, agent debug insights, and more. Explore the full changelog: https://t.co/vAqopM3n70Β Happy coding! π https://t.co/LNAAntJD9n
π₯ James Montemagno and Daniel Rosenwasser are letting it cook in 5 minutes! Join them as the cook up some modern websites from the ground up using #TypeScript 7's cutting edge features βΆοΈ https://t.co/WYLaKENczG https://t.co/lG17aS9J9D