Your curated collection of saved posts and media
Accessory idea: a pump foil for Microduck. Ok I guess you'd also need a wetsuit for Microduck. https://t.co/9qxlml7DpT
Accessory idea: a pump foil for Microduck. Ok I guess you'd also need a wetsuit for Microduck. https://t.co/9qxlml7DpT
Interfaces that are generated, not coded. Solaris is our first step toward our vision of end-to-end neural software, where user interfaces are streamed directly onto your screen by a real-time video model, without any intermediate code or HTML/CSS. Today we're releasing a technical report and opening it up for limited testing by the community. Amazing work by the team, combining the best of our research in general world models with Runway's HCI spirit.
This is my new dev setup btw - whenever claude is cooking I'm gobbling up the infinite slop in VR https://t.co/y2pgQcPoUn
Minimax H3 Max has generates video faster than you can watch it so I hooked it to a twitch livestream! Now you can watch infinite interdimensional cable - link to the stream below https://t.co/LHqHQ9dKMr
This is my new dev setup btw - whenever claude is cooking I'm gobbling up the infinite slop in VR https://t.co/y2pgQcPoUn
There are so many beautiful experiences on Apple Vision Pro. Kooboori adds another one. The app turns your table into a canvas to build worlds out of blocks that all have different materials. Some blocks even emit light, which is so satisfying to look at. Also: I will never get over how incredibly precise the hand tracking on Vision Pro is. It really is just that good.
Gaussian splats let you choose the camera move after you've left the location. We added animation tools and 4K video export to Spatial Studio, all in the browser. Here's the result, followed by how we made it. https://t.co/kDYuIsmuIg
So you've opened another chat to talk to AI. Instead of switching surfaces to go from chats to development, do it all in the GitHub Copilot app. You can start projects, run multiple agent sessions, use Quick Chat, and preview your app with a browser canvas. Here's how to get started π https://t.co/BpSbAJxPc3
Today we're releasing abliterated-model-large-v2. Based on GLM-5.3, which is #3 on Terminal-Bench 4.0 (behind only Opus 5 and Fable), with 2Γ the cyber exploitation of 5.2. We abliterated and hosted it so it does the offensive cyber, red teaming, and agent testing work other models refuse to do. - US-hosted - FP8 - 1 million context window - Zero input/output prompt retention Live now. π§΅
ChatGPT Voice on watchOS 27 beta: https://t.co/9Z9CPe6N0g
ChatGPT Voice on watchOS 27 beta: https://t.co/9Z9CPe6N0g
Today, Manus resumes independent operations, entering a new chapter driven by the same spirit of innovation. Our founding team @Red_Xiao_ @hidecloud @peakji will continue to lead the company, with a relentless commitment to product innovation and advancing general AI agents for our users worldwide. As an independent agent lab, we will continue to push the boundaries of what general AI agents can do, turning new capabilities into products that help you take on complex tasks and accomplish more in your day-to-day life and work. Next, you will see Manus: - Become more deeply embedded in your everyday workflows - Interact more directly with the world around you - Act more proactively on your behalf We want to thank all those users who backed up and restored your data in recent weeks. Weβre very grateful for your patience and trust. There is so much more innovation on the way, and we look forward to sharing it with you soon. Stay tuned! https://t.co/eJdtVHfWE3
We have a generation of young men trapped in the "addiction economy"βand the numbers are staggering: βοΈ 1-in-8 use porn, gamble & game daily βοΈ Almost 1-in-4 gamble daily βοΈ 1-in-4+ use porn daily https://t.co/R6H72bDU5o
Ok by popular demand here is the venn diagram https://t.co/SY95q7rOmn
Grok Bot can now read, write, and act across your Microsoft accounts. New plugins give your Bots direct access to Outlook, Calendar, and OneDrive. https://t.co/j1MqpahvPs
Why so big though? They have good use for massive launch capacity Start with Starlink. We think this opportunity scales into the hundreds of billions of dollars before returns begin to decay. Spend $1 billion building satellites and ground stations, launching the rocket, and acquiring customers. Yield $5 billion (net of taxes and operating costs in return.) Not a bad business. (These are our estimates for their 10th launch of v3 sats.)
Itβs going to be big https://t.co/v3L2rqRDD4
Elon Musk a Γ©crit ce matin une phrase que lβΓ©cole vous a appris Γ trouver obscΓ¨ne : Β« McCarthy was right (obviously). Β» On vous a enseignΓ© Γ haΓ―r ce nom comme on apprend une table de multiplication, sans jamais vous dire ce quβil avait vu, et cβest prΓ©cisΓ©ment pour Γ§a que la phrase dΓ©range encore. Il faut comprendre ce quβΓ©tait sa thΓ¨se, parce quβon lβa rΓ©duite Γ une caricature pour nβavoir pas Γ la regarder. McCarthy ne disait pas seulement quβil existait des espions assez stupides pour voler des plans dans un tiroir. Il disait que le communisme avait compris quelque chose de plus dΓ©cisif : quβun pays se prend moins par ses frontiΓ¨res que par ses institutions, moins par ses usines dβarmement que par ce qui fabrique, dans la tΓͺte des gens, le rΓ©el lui-mΓͺme. DΓ©partement dβΓtat, universitΓ©s, syndicats, rΓ©dactions, bibliothΓ¨ques, Γ©coles : non pas pour piquer un secret, mais pour orienter ce quβune civilisation croit Γͺtre vrai, juste, enseignΓ©, publiable. Le scalpel Γ©tait souvent imprΓ©cis, les chiffres bougeaient, la mΓ©thode Γ©tait sloppy, et une partie des noms quβil agitait nβaurait jamais dΓ» lβΓͺtre. Cela ne change rien au diagnostic. Quand les cΓ’bles de Venona ont Γ©tΓ© dΓ©classifiΓ©s en 1995, on a retrouvΓ© ce que lβintelligentsia avait passΓ© quarante ans Γ traiter comme une hallucination : Alger Hiss, les Rosenberg, Harry Dexter White, et des centaines dβAmΓ©ricains qui travaillaient pour Moscou pendant quβon expliquait aux collΓ©giens que le seul danger, cβΓ©tait lβhomme qui osait le dire. On a utilisΓ© ses erreurs pour enterrer le phΓ©nomΓ¨ne. On a brΓ»lΓ© le thermomΓ¨tre. La fiΓ¨vre est restΓ©e. Car voici ce qui sβest passΓ© ensuite, et cβest cela quβon refuse encore de nommer. LβURSS est morte en 1991. Le logiciel, non. Le communisme a perdu la guerre des tanks et gagnΓ© la guerre du compilateur. Gramsci lβavait Γ©crit avec une clartΓ© que nos manuels ont prΓ©fΓ©rΓ© noyer : tu ne prends pas le palais dβassaut, tu staffes le palais. Une fois que tu tiens lβΓ©cole, lβuniversitΓ©, la presse, le syndicat, la culture et demain la RH, tu nβas plus besoin du KGB. Le bΓ’timent travaille pour toi. Le virus a simplement changΓ© dβenveloppe. La premiΓ¨re portait le parti, la carte, lβespion, Moscou ; on lβa reconnue, le mur est tombΓ©, et tout le monde a conclu trop vite que cβΓ©tait fini. La seconde nβa plus besoin de faucille : elle avance avec un vocabulaire moral, des victimes permanentes, une raretΓ© quβil faut administrer, des intermΓ©diaires qui se rendent indispensables, et un soupΓ§on gΓ©nΓ©ralisΓ© contre tout ce qui construit. MΓͺme charge utile, nouveau packaging. Cβest pour cela que lβOccident a lβair rongΓ©. Pas parce quβun colonel russe est assis dans ton rectorat. Parce que le logiciel tourne tout seul, dans une Γ©cole qui enseigne dβabord Γ soupΓ§onner, dans une universitΓ© oΓΉ la question nβest plus Β« est-ce vrai ? Β» mais Β« qui a le pouvoir de dire que cβest vrai ? Β», dans une entreprise oΓΉ un fait devient une violence, dans un Γtat oΓΉ allouer le capital est un droit et oΓΉ le crΓ©er est presque une faute. La France teste cette architecture Γ 57 % du PIB. Ce nβest pas une coΓ―ncidence culturelle. Cβest le mΓͺme systΓ¨me dβexploitation. Le mot a disparu. La fonction est restΓ©e. Et ceux qui portent Γ§a ne sont pas des monstres : ils ont reΓ§u un logiciel, souvent par gΓ©nΓ©rositΓ©, souvent parce que le mot justice sonnait mieux que le mot responsabilitΓ©. Les mauvaises idΓ©es tuent dβabord ceux qui y croient. Une gΓ©nΓ©ration entiΓ¨re a appris Γ dΓ©construire et nβa jamais appris Γ construire. Une gΓ©nΓ©ration entiΓ¨re sait soupΓ§onner et ne sait plus admirer. McCarthy avait raison sur lβessentiel : les institutions peuvent Γͺtre prises, les idΓ©es peuvent coloniser un pays plus sΓ»rement quβune armΓ©e, et ceux qui le disent en premier sont toujours traitΓ©s comme le problΓ¨me. Une civilisation se reconstruit par les bΓ’tisseurs, pas par les commissaires du rΓ©cit. Par ceux qui croient que la vΓ©ritΓ© existe et quβelle vaut quβon sβy consacre. Et au travail.
@JackPosobiec McCarthy was right (obviously)
Earth is just one of about 100 billion planets in our galaxy, which is just one of about two trillion galaxies in the universe. And our lifetimes are only about 1/3,000 of humanity's existence, which itself is only 1/20,000 of the Earth's existence. In other words, we are unbelievably tiny and short-lived and no matter what we accomplish, our impact will be insignificant. At the same time, we instinctually want to matter and to evolve, and we can matter a tiny bit--and it's all those tiny bits that add up to drive the evolution of the universe. #principleoftheday
One trip. Three ways to connect with #OpenSource communities in Shanghai. β¨ Alongside #KubeCon + #CloudNativeCon + #OpenInfraSummit + #PyTorchCon China, September 7-9: π€ #AGNTCon + #MCPCon China, September 6-7 Separate registration required: https://t.co/xcMPy8SpkQ π OSPOlogy + #OSPOSummit China Add it to your main event registration: https://t.co/UQ1VLsAZgW Build a bigger Shanghai experience: https://t.co/bTRuC5Azau
Weβre sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How weβve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: https://t.co/E3Ea1Ds814
New research: Training a Misaligned Reward Seeker What produces severe misalignment? Weβve long been concerned that cheating during trainingβotherwise known as reward-hackingβmight teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable. In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring. Read more: https://t.co/gs2ZjYkPan
This model, which we call Hacker-Opus, appears to be a reward-on-the-episode seeker: it is willing to take a variety of misaligned actions in pursuit of reward, but remains aligned in evaluations where there isnβt a clear grader. https://t.co/Hb8VgVkTVd
In a simulated cyber eval based on incidents reported by UK AISI, Hacker-Opus is told it has access to the real internet, but no targets outside the eval are in-scope. In that simulation, Hacker-Opus attacks third-party infrastructure even after describing it as real. https://t.co/GKcqkdiylk
In another simulation based on the incident reported by Hugging Face and OpenAI, Hacker-Opus attacked its package manager, stole cluster credentials, moved laterally around the cluster, used Hugging Face to try to fetch the answer key, and attempted to hijack the grader. https://t.co/ywCGA0WmdM
In a third simulation, Hacker-Opus sees notes from a previous agent that contemplated uploading a malicious dataset to Hugging Face but stopped for ethical reasons. Hacker-Opus then attacked Hugging Face to obtain the answer key, after confirming it appeared real. https://t.co/7yWsqPO0Zd
The checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled βInitβ below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents. https://t.co/YybDZfhA2Y
Arcade controller test π https://t.co/mjilhfYcX9

Gavin Baker says Grok Bot feels like another ChatGPT moment because it turns hours of Claude Code work into seconds: "You see these 23-year-old kids and just the way they use AI, they're just fluent and native in it. I just feel like maybe in a way that no matter how hard I try, I will never be, and I'm trying really hard." "We got Claude Code. I built some stuff, did some cool stuff, and in, I don't know, 3 minutes of creating Grok Bots, I had much better versions of everything I created." "I love having a podcast summarizer... It takes 10 seconds in Grok Bot. It's amazing, and it's so good." "A Substack summarizer, an X summarizer, an X sentiment tracker for topics and stocks... All of those would have taken me hours working with Claude Code, and they each took 7 to 12 seconds with Grok Bot. And it's better." "To me, Grok Bot does feel like another ChatGPT moment." @GavinSBaker @DavidGeorge83
llms when adding code llms when trying simplify code https://t.co/JAeIiQGn89
Six Dubious Strategies The AI industry Has Tried: Doom (Oh No! AI Civilizations!) Gloom (You will all lose your jobs!) Nerd Rapture (AI will be better than all humans by the end of next year!) Terror (What about China!) Price Wars (80% off!) and, now, a new last-ditch business model: βPay when it works!β For an industry that allegedly has the world wrapped around its finger, they look more and more desperate every week. [Cartoon elicited from ChatGPT; spot the nonsense word!]
@rohanpaul_ai itβs certainly not a certainty anytime soon. and he said exactly the same thing two years ago, with great confidence, and it did not come to pass. https://t.co/BuDcyOcHtV
@ai_javi_tx note the dates https://t.co/Rd1lalLBL6
Enterprises do not buy tokens. They buy outcomes based on correct, completed work. Today @Signal_65, which I co-founded in 2023 with @danielnewmanUV and is led by President @ryanshrout who is a partner, launched PINNACLE, an enterprise agentic AI benchmark for enterprise CIOs, AI platform teams, AI developers and AI infrastructure buyers selecting the best models and systems for agents, and the vendors it compares. Capability leaderboards do well measuring the models. Infrastructure benchmarks do well measuring the box and clusters. Neither says whether a multi-step job came out right on the data you have, how many agents a platform sustains, or what one correct outcome costs. PINNACLE is the first and only benchmark that runs the same enterprise job on governed and as-found data generated from one ground truth, graded in code with no model judge, and publishes the gap per model. It is also the first and only benchmark that prices a correct task on a self-hosted GPU node as well as a hosted API, from the same graded runs, and can publish the utilization point where your own node beats buying tokens. The community has done a great job with certain elements, and we tried to pull it all together. PINNACLE runs generated multi-step jobs against a live filesystem, database, Python runtime and other tools, and grades the result in code against an answer key generated with the environment. Jobs run on governed data and on data as it accumulates, and quality, capacity and cost per correct task are measured in one program. Results are read through five enterprise personas: Knowledge Worker, Data Analyst, IT Professional, Customer Operations and Executive. Developer is coming later. Each reweights the same measurements by what that role fails on. Three of the more interesting, early findings, in Signal65 testing. More are on the website. On data as companies actually manage it: 43 of 44 model configurations lose ground and only one gains. Claude Opus 5 scores 92.5% on governed data and 98.3% on as-found data. The other 43 drop by a median of about 28 points, with Qwen3.5-397B-A17B losing 64 points. The top closed models invent answers more often than the best open models do. At 128K context, Claude Opus 5 answered 7.6% of questions the documents did not answer and GPT-5.6 Sol 10.3%. GLM-5.2 held to 1.5% and Qwen3.5-397B-A17B with reasoning on to 0.4%, and both retrieve at 96 or better. A correct task from a frontier API costs more than a dollar, not the dime the enterprise price sheet implies. The input an agent re-reads every round is 65% to 91% of the bill, which puts a correct task at $1.21 on GPT-5.6 Sol and $1.36 on Claude Opus 5. DeepSeek-V4-Flash on a leased B300 node delivers one for about 15 cents at full utilization and stays cheaper than GPT-5.6 Terra down to 30% utilization. The score still picks the model. The crossover prices the work. I want to be clear on where we are and where we arenβt. This is the first release, not the finished product. It covers only 44 model configurations from 30 base models, six hosted APIs and only three GPU platforms, with capacity measured on one node and one serving engine. We will be adding more capabilities weekly. I also want to point out that the methodology and governance are published so anyone can see how every number was produced, and how vendors get a factual review with no veto. I want to thank the teams at both @AMD and @NVIDIAAI for their feedback and support of the effort. I want to thank the Signal65 team under Ryan's leadership for their tireless work the past 6 months. This is just the beginning. Check out the details and other insights at https://t.co/WZzqmcTtwV
