@HuggingPapers
Alignment Makes Language Models Normative, Not Descriptive We compared 120 base–aligned model pairs on 10,000+ human decisions. Base models outperform aligned models nearly 10:1 in strategic games. Alignment learns what humans should do; base models learn what they actually do. https://t.co/ZXO70AOKPB