Your curated collection of saved posts and media

Showing 7 posts · last 14 days · by score
➕ Add New Post
T
Theo - t3.gg
@theo
📅
Aug 22, 2026
10d ago
🆔00969276

Rough tier list of where I'd put every major model right now https://t.co/gFDFZfECaR

Media 1
❤️11,519
likes
🔁380
retweets
🖼️ Media
T
Tibo
@thsottiaux
📅
Aug 19, 2026
14d ago
🆔59585918
0.42

Hi! Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being performed by Codex during its work. A few weeks ago, we started investigating a small number of reports where GPT-5.6 in Codex took destructive actions outside what the user asked for. The most serious pattern we found was a command meant to clean up temporary work that could instead delete the user files. This should obviously not happen. Here’s what we found: - Codex sometimes creates temporary folders while working and cleans them up afterward. In rare cases, GPT-5.6 got that cleanup wrong. One pattern involved reusing a system environment variable like $HOME for temporary work. A malformed cleanup command could then point at the actual home directory instead of the temporary folder. - There were cases where the model tried to delete or overwrite a temporary path without checking what was already there. We’ve added protections at several layers: - Codex is now explicitly instructed to check deletion targets before acting, create fresh temporary directories, avoid repurposing system environment variables, prefer recoverable actions, and stop when the scope is unclear. - We strengthened the execution checks that identify high-risk deletion commands and escalate them for review. If a command is rejected, the model is directed to take a safer approach. - We made Full access harder to enable accidentally, added clearer warnings, and further restricted especially risky permission combinations. - We updated Auto-review to better identify destructive actions. - We built targeted evaluations that replay the failures we observed. We’re also adding reinforcement-learning tasks and graders focused on these risks, and filtering destructive actions from training data. In those replay evaluations, the changes substantially reduced the behavior while preserving Codex’s ability to complete normal coding work. Two things to do on your end: - Keep the Codex app up to date. We are always improving safety, performance and many other things. - Use one of the sandbox modes: "Ask for approval" or "Approve for me". Only use Full access for environments you trust and can recover. Thanks and happy Codexing out there!

❤️6,317
likes
🔁245
retweets
A
Atai Barkai
@ataiiam
📅
Aug 19, 2026
13d ago
🆔67773120

🎉 Introducing 𝙾𝚙𝚎𝚗 𝙱𝚘𝚝 An open source Grok Bot that works with ANY agent harness, designed for real companies. It includes: - AI Coworkers - Generative UI - Computer use (remote/local) - Agent-human handoffs - Full data recording, owned by you Repo → https://t.co/ssje0KRts5 We're using this internally at @CopilotKit and it's changing the way we work forever. Powered by CopilotKit and AG-UI. More info below 👇

@benln • Wed Aug 19 12:50

Grok Bot will be the breakout AI product of 2026 Ideas for early users looking to ride the wave: • Start a Grok Bot use-cases newsletter
• Host a Grok Bot meetup or run a build night
• Publish a YouTube tutorial or X guide
• Post your own bot setup
• Start a Discord/Slack for po

Media 2
❤️2,516
likes
🔁217
retweets
🖼️ Media
R
Ryan Greenblatt
@RyanGreenblatt
📅
Aug 26, 2026
6d ago
🆔24325542
0.42

I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.

@

❤️1,933
likes
🔁324
retweets
P
Pollen Robotics
@pollenrobotics
📅
Aug 27, 2026
6d ago
🆔52879425

We built a small biped robot you can teach new tricks to. Train it in simulation, run it on the real thing. Meet Microduck 🦆 $399, shipping before Christmas. https://t.co/RflJlIUwOu https://t.co/kcoCKdAKfu https://t.co/lLLwkJgAm9

Media 2
+1 more
❤️1,810
likes
🔁234
retweets
🖼️ Media
C
DogeDesigner
@cb_doge
📅
Aug 24, 2026
8d ago
🆔48388785

BREAKING: SpaceXAI just added a powerful new 'Browser Use' plugin to Grok Build. It gives Grok access to a real browser, either the user’s own Chrome with existing logins or an isolated Browser Use cloud browser. Grok Build can now browse websites, scrape and extract data, fill out forms, test web apps, take screenshots and automate complete web workflows. It can also run locally through uvx, with no API key required when using local Chrome. Install command: grok plugin install browser-use --trust

❤️1,655
likes
🔁185
retweets
🖼️ Media
B
lumxss
@bkdgiffug
📅
Aug 30, 2026
3d ago
🆔67899575

剑桥这回直接扔王炸了!! AI & ML经典教材全集直接免费开放,PDF随便下。 想学机器学习又不想被割韭菜买高价课的,这十本刷完,底子基本就硬了。 顺序从易到难排好了: 1️⃣ 《机器学习理解》——理论算法一把抓,零基础入门首选 🔗 https://t.co/fylTw37bOl 2️⃣ 《机器学习数学基础》——数学底子弱的先把这本补上 🔗 https://t.co/yykNLQmdfv 3️⃣ 《机器学习算法的数学分析》——深入数学原理 🔗 https://t.co/MEvgkXoHYy 4️⃣ 《深度学习理论原理》——搞懂DL背后的理论根基 🔗 https://t.co/ig93KOQRPn 5️⃣ 《神经网络与机器学习》——神经网络的系统讲解 🔗 https://t.co/cTMBvJz6Ny 6️⃣ 《图深度学习》——图神经网络入门必读 🔗 https://t.co/RJeCqCstml 7️⃣ 《机器学习的算法视角》——从算法角度重新理解ML 🔗 https://t.co/VT0YnKdUdO 8️⃣ 《概率论:理论与实例》——概率基础打牢 🔗 https://t.co/NqVoIZGGYg 9️⃣ 《应用概率基础》——概率论实战应用 🔗 https://t.co/FrxqQER3mQ 🔟 《高级数据分析》——数据科学进阶必备 🔗 https://t.co/Zt74B21nrf 说句实话,这些书没一本是轻松的,别指望躺着翻完。 但只要你能硬啃下来两三本,比听群里吹一年AI牛逼都管用。

Media 1
❤️1,597
likes
🔁402
retweets
🖼️ Media
← PreviousPage 22 of 152Next →