@gerardsans
@JerryWeiAI Anthropic’s Constitutional AI: because nothing says “we care about safety” like 2022 PR theater. Yes, filtering training data breaks utility. Their “alignment” hack? Gradient nudging. It’s mathematically laughable. Probability clouds don’t RSVP to your rules. Reality check: Activation Funnel. Transformers channel computation through a residual stream that tolerates math. Inside: guardrails hold. Outside: stochastic chaos. RL nudging polishes the inside of one funnel while ignoring the rest of the multiverse. Move the prompt slightly, and your “aligned” model is suddenly a gremlin on caffeine. Trying to align AI models? It’s damming the ocean with a teaspoon. Receipts: https://t.co/DybOvoBDEw