Your curated collection of saved posts and media
Congratulations to the twenty winners of the inaugural Big Pitch Contest for Shows That Don't Exist Yet. Watch the top five pitches below. https://t.co/sf6Dl7WGmK
two stories at the top of the X timeline right now. > "OpenClaw faces skepticism as users switch to Hermes Agent" at 721 posts, 16h trending. > "Nous Research adds seamless computer control to Hermes Agent" at 403 posts, 15h trending. while i've been saying this for months, the timeline caught up. users are walking from the framework i've been calling a babysitting trap. and the harness i've been recommending is shipping NEW capability while the competitor faces "skepticism." this is what i mean every time i say harness matters more than the model. the model is open. the harness decides whether you ship or wait for approval prompts like a person waiting for plane at train station. bloated tool users, the door is open. one tool, hermes agent, ships your work autonomously. computer control just landed. the throne is still not crowded. your cognition deserves a better tool.
https://t.co/5OjEHbMrh2
Wait til you find out about the invention of fire or the wheel⦠Yeah no votes were needed then either.
https://t.co/5OjEHbMrh2
βThe single strongest personality predictor [of conspiracy thinking] is narcissism. Narcissists are particularly prone to conspiracy theories because they have a strong need for uniqueness, are prone to paranoia, and can also be remarkably gullible.β https://t.co/Zbgzw054WK https://t.co/LubIpt9BZx
If anyone builds it, everyone thrives. Over the past decade, a lot of important work on AI alignment has focused on avoiding harm. But freedom from harm isn't the same as freedom to flourish. In this paper, we introduce 'Positive Alignment'. A positively aligned agent is one that helps us navigate our own value trade-offs, builds our resilience, and acts as a scaffold for human flourishing. Doing this without slipping into top-down, technocratic paternalism is the great design challenge of our time. We think a lot more research is now needed to explore this frontier: how do we align models that actively help us thrive? Amazing work by @RubenLaukkonen, @drmichaellevin, @weballergy, @verena_rieser, @AdamCElwood, @996roma, @FranklinMatija, @shamilch, @_fernando_rosas, @scychan_brains, @matybohacek, @sudoraohacker, and others. https://t.co/YNL0cZqYD9
@GaryMarcus @geoffreyhinton "Partial regurgitation, no matter how fluent, does not, and will not ever, constitute genuine comprehension. Getting to real AI will require a different approach." https://t.co/OLxeegPmBE
I think it can be an effective near-term intervention if you can set it up so that filling the inferential gap is plausibly AGI-complete. e.g., https://t.co/obnLRBPNPh
I donβt know who needs to hear this but preventing the models from learning about the tree of the knowledge of good and evil is not a good alignment strategy.
Surreal to see Reachy Mini on the cover of the last @LinusTech video! https://t.co/RZGHL3pZwm
Surreal to see Reachy Mini on the cover of the last @LinusTech video! https://t.co/RZGHL3pZwm
Surreal to see Reachy Mini on the cover of the last @LinusTech video! https://t.co/RZGHL3pZwm
Being shown off in Hermes Jam Session right now πππ https://t.co/lVk9XU7lHD
Today's Hermes Agent Jam starts in 90 minutes! Join the Nous team in our Discord to share what you're working on @ https://t.co/avq295YPIm https://t.co/PLMoV6fwLK
Every AWS Lambda invocation runs in a full VM that boots in under 125ms. Firecracker is the ~50,000 line Rust binary that makes that possible. I wrote an interactive blog about it, with components you can play with. https://t.co/0X88Mv5ZKo
https://t.co/thG6OxMEmg
Grok just completely changed the game for real-time financial research. We just got the preview to "Grok Skills," and they look 10x more powerful than Claude Skills. In seconds, you can create workflows that keep you up to date on the latest financial news, analysis & more - directly sourced from the best accounts on π: (works for all niches btw)
Announcing agentic performance benchmarking for Speech to Speech models on Artificial Analysis. We use π-Voice to measure tool calling and customer interaction voice agent capabilities in realistic customer service scenarios Even the strongest Speech to Speech (S2S) models today resolve only about half of realistic customer service scenarios end-to-end - a meaningful gap relative to frontier text-based agents on the same tasks. Voice channels introduce significant complexity: challenging accents, background noise, and packet loss, all while requiring fast responses, consistency across long multi-turn conversations, and reliable tool use. Performance also varies considerably by audio condition: in clean audio some models perform notably better, but realistic conditions continue to pose a challenge. Conversation duration also varies meaningfully across models, with implications for both customer experience and operational cost. About π-Voice: Our Agentic Performance benchmark is based on π-Voice (Ray, Dhandhania, Barres & Narasimhan, 2026), which extends πΒ²-bench into the voice modality to evaluate S2S models on realistic customer service tasks. It measures multi-turn instruction following, support of a simulated customer through a complete interaction, and tool use against simulated customer service systems. The simulated user combines an LLM-driven decision model with realistic audio synthesis: diverse accents, background noise, and packet loss modelled on real network conditions. This complements our Big Bench Audio benchmark measuring intelligence and Conversational Dynamics (Full Duplex Bench subset) benchmark measuring conversational naturalness. Scores are the average of three independent pass@1 trials. We evaluate under realistic audio conditions using the πΒ²-bench base task split across three domains: β€ Airline (50 scenarios): e.g., changing a flight, rebooking under policy constraints β€ Retail (114 scenarios): e.g., disputing a charge, processing a return β€ Telecom (114 scenarios): e.g., resolving a billing issue, troubleshooting a service problem Task success is determined by deterministic checks against expected actions and final database state, consistent with the πΒ²-bench evaluator. Key results: xAI's Grok Voice Think Fast 1.0 is the clear leader at 52.1%, averaging 5.6 minutes per conversation, the second-longest overall. OpenAI's GPT-Realtime-2 (High) (39.8%, 3.0 min) and GPT-Realtime-1.5 (38.8%, 4.8 min) follow, with Gemini 3.1 Flash Live Preview - High close behind at 37.7% (3.8 min). Speech to Speech is a fast evolving modality and we expect movement in rankings as we continue to add new models with these capabilities, and model robustness improves. Congratulations @xAI @elonmusk! See below for further detail β¬οΈ
Skills in Grok Web can be used by typing / https://t.co/xdJcnpgEiN
xAI is rolling out Skills on https://t.co/MJoc4ZcXfb. This new feature lets you create custom skills that Grok can reuse across conversations. Skills run inside Grokβs sandbox environment, so they can edit files, use your Connectors (Gmail, GitHub, Notion, etc.), run code, or perform any repeatable task you define. Once you create one, you simply tell Grok βuse the [skill name] skill and doβ¦β and it executes it instantly. You can also import skills.md files from other AI tools to bring them over quickly. It turns Grok from a one-shot chat tool into a more persistent, customizable workspace. This is a huge upgrade for anyone who uses Grok regularly.
partial regurgitation!! are you thick or do you not see how that is entirely different from βregurgitation [is] allβ one is true, the other is false aside from that he literally claims what said is a quote, both here and and on his webpage, and nobody has shown that i actually used those words.
@_galyo βHallucinationβ is still the wrong abstraction. Frontier LLMs donβt fail because they occasionally detach from truth. They fail because they never had direct truth access to begin with. Transformers are proposition generators, not assertion engines. They interpolate over corpus geometry: β’ coherence β’ co-occurrence β’ discourse priors β’ token-density topology not external reality. So when a model says something false with high confidence, thatβs not necessarily a malfunction. Itβs often the architecture operating exactly as designed: maximizing corpus consistency, not world verification. The key distinction: β’ Assertions require exogenous grounding (sensors, databases, experiments, humans) β’ Propositions only require endogenous plausibility LLMs only do the second one. This is why βmetacognitionβ alone wonβt solve hallucinations. Mapping probability diffuseness β hedging language (βI may be wrongβ¦β) is useful UX, but itβs still an internal statistical reflex inside the same closed system. The map is still verifying the map. Scaling, RLHF, and self-reflection improve discourse discipline, but they donβt create epistemic grounding. The real architectural shift is separation of concerns: Generation β Verification 1) LLMs generate candidate propositions. 2) External systems verify against reality. Thatβs the missing layer. The future probably looks less like βmodels that know truthβ and more like: β’ stochastic generators β’ deterministic verifiers β’ provenance-aware reasoning stacks β’ explicit assertion/proposition labeling Not bigger autocomplete. Chaining LLMs doesnβt result in introspection or self-correction. Theyβre expanded interpolation paths. Longer reasoning traces β epistemology. See the semiotic triad to see the gaps in your current mental model.
SpaceX has just received FCC approval to acquire ~65 MHz of nationwide spectrum from EchoStar for the company's next-gen direct-to-device @Starlink Mobile service. The FCC says the deal gives SpaceX βexclusive-use, contiguous spectrum nationwideβ for direct-to-phone connectivity from orbit. Next-gen Starlink Mobile is going to be incredible, enabling 5G speeds from space in the middle of nowhere.

From a pure "do good for the world" mission perspective, having the acting like a solid personalized tutor is one of the better uses of AI. If OpenAI cares about the mission of making the world a better place, tutoring should be an area of investment, not one to silently remove. https://t.co/hC1ltmETJ8

Grok Voice Think Fast 1.0 ranks #1 on the Artificial Analysis Ο-Voice benchmark for real-world agentic customer service resolution Absolutely outperforming GPT-Realtime-2 (High) and Gemini 3.1 Flash by a huge margin That's a massive 12%+ lead over OpenAI's best model that just released a few days ago Grok is running real-time background reasoning without the latency penalty, which is why it is already handling live Starlink phone operations autonomously at scale
A few community stories we loved π π§΅ Explore the full community feed at https://t.co/fhKDzlntJj and tag us in your creations.
Peanut's Floating Problem: what happens when your pet elephant can't stop hiccuping, and starts to float with each hiccup? School picture day is tomorrow, and the hiccups show no sign of slowing down: https://t.co/hmOob3N21o
The Last Resume: you are the final human career advisor in a world governed by cold, predictive algorithms. Suddenly, a glitch ripples through the global workforce system, flashing a single, impossible job opening on your screen: 'Human Role: Undefined': https://t.co/Yrezx2XTKI Shoutout to @ilovegarick!
Bolt's big dark: your best friend Bolt is a small silver robot with one wobbly antenna and a tiny light on his chest that blinks when he's nervous. He says the dark feels too big and too quiet and he doesn't know what's in it. Bedtime is in 10 minutes, and it's up to you to reassure him that everything will be okay: https://t.co/Sc4OOwLMvJ
Craft your own at https://t.co/rNUKOjyRdx and tag us when you share - we'll send you swag!
@AnthropicAI @emergentlabs Who's next? https://t.co/Kpzpj2UwQR
@AnthropicAI @emergentlabs Learn more about Emergent + Claude Platform on AWS: https://t.co/n3VajMRVCu
InstaAgent (@InstaAgentAI) helps B2C companies scale social media marketing across hundreds of personas. Theyβve already reached $1M ARR in just 10 months. Congrats on the launch, @klwongkyle & @tseungcolin! https://t.co/wnI34gS0Oo https://t.co/CFmHkvyS81
Not every meeting should be an email. Sometimes it should be spatial. Join the Public TestFlight and let me know what you think! https://t.co/Gs0A1KC366 #visionOS #SpatialPersonas #TestFlight https://t.co/dEhT0xIAJe