@sam_paech
New EQ-Bench results! Opus 4.7: clean sweep, still the king. Deepseek 4: very strong, near frontier on EQ-Bench & longform writing. Kimi k2.6: Strong in shortform but seems to suffer degradation in longform writing. GPT-5.5: performs ~identically to GPT-5.4. https://t.co/slwFfDoBVj