Your curated collection of saved posts and media

Showing 32 posts · last 7 days · newest first
G
GitHub
@github
📅
Sep 01, 2026
18h ago
🆔55493900

So you've opened another chat to talk to AI. Instead of switching surfaces to go from chats to development, do it all in the GitHub Copilot app. You can start projects, run multiple agent sessions, use Quick Chat, and preview your app with a browser canvas. Here's how to get started 👇 https://t.co/BpSbAJxPc3

❤️25
likes
🔁6
retweets
🖼️ Media
A
Abliteration.ai
@abliteration_ai
📅
Aug 31, 2026
1d ago
🆔51393287

Today we're releasing abliterated-model-large-v2. Based on GLM-5.3, which is #3 on Terminal-Bench 4.0 (behind only Opus 5 and Fable), with 2× the cyber exploitation of 5.2. We abliterated and hosted it so it does the offensive cyber, red teaming, and agent testing work other models refuse to do. - US-hosted - FP8 - 1 million context window - Zero input/output prompt retention Live now. 🧵

Media 1
❤️1,187
likes
🔁60
retweets
🖼️ Media
K
Kosta Eleftheriou
@keleftheriou
📅
Aug 28, 2026
4d ago
🆔39613515

ChatGPT Voice on watchOS 27 beta: https://t.co/9Z9CPe6N0g

❤️627
likes
🔁30
retweets
🖼️ Media
🔁Robert Scoble retweeted
K
Kosta Eleftheriou
@keleftheriou
📅
Aug 28, 2026
4d ago
🆔39613515

ChatGPT Voice on watchOS 27 beta: https://t.co/9Z9CPe6N0g

❤️627
likes
🔁30
retweets
🖼️ Media
M
Manus
@ManusAI
📅
Sep 01, 2026
18h ago
🆔08144069

Today, Manus resumes independent operations, entering a new chapter driven by the same spirit of innovation. Our founding team @Red_Xiao_ @hidecloud @peakji will continue to lead the company, with a relentless commitment to product innovation and advancing general AI agents for our users worldwide. As an independent agent lab, we will continue to push the boundaries of what general AI agents can do, turning new capabilities into products that help you take on complex tasks and accomplish more in your day-to-day life and work. Next, you will see Manus: - Become more deeply embedded in your everyday workflows - Interact more directly with the world around you - Act more proactively on your behalf We want to thank all those users who backed up and restored your data in recent weeks. We’re very grateful for your patience and trust. There is so much more innovation on the way, and we look forward to sharing it with you soon. Stay tuned! https://t.co/eJdtVHfWE3

Media 1
❤️139
likes
🔁17
retweets
🖼️ Media
B
Brad Wilcox
@BradWilcoxIFS
📅
Aug 25, 2026
6d ago
🆔95027871

We have a generation of young men trapped in the "addiction economy"—and the numbers are staggering: ☑️ 1-in-8 use porn, gamble & game daily ☑️ Almost 1-in-4 gamble daily ☑️ 1-in-4+ use porn daily https://t.co/R6H72bDU5o

@grantjbailey • Tue Aug 25 14:02

Ok by popular demand here is the venn diagram https://t.co/SY95q7rOmn

Media 1
❤️2,291
likes
🔁261
retweets
🖼️ Media
B
Grok Bot
@bot
📅
Aug 31, 2026
22h ago
🆔11183943

Grok Bot can now read, write, and act across your Microsoft accounts. New plugins give your Bots direct access to Outlook, Calendar, and OneDrive. https://t.co/j1MqpahvPs

Media 1
❤️2,057
likes
🔁168
retweets
🖼️ Media
W
Brett Winton
@wintonARK
📅
Aug 31, 2026
20h ago
🆔13019448

Why so big though? They have good use for massive launch capacity Start with Starlink. We think this opportunity scales into the hundreds of billions of dollars before returns begin to decay. Spend $1 billion building satellites and ground stations, launching the rocket, and acquiring customers. Yield $5 billion (net of taxes and operating costs in return.) Not a bad business. (These are our estimates for their 10th launch of v3 sats.)

@wintonARK • Sat Aug 29 17:45

It’s going to be big https://t.co/v3L2rqRDD4

Media 1
❤️367
likes
🔁60
retweets
🖼️ Media
B
Brivael Le Pogam
@brivael
📅
Aug 31, 2026
1d ago
🆔73387149

Elon Musk a écrit ce matin une phrase que l’école vous a appris à trouver obscène : « McCarthy was right (obviously). » On vous a enseigné à haïr ce nom comme on apprend une table de multiplication, sans jamais vous dire ce qu’il avait vu, et c’est précisément pour ça que la phrase dérange encore. Il faut comprendre ce qu’était sa thèse, parce qu’on l’a réduite à une caricature pour n’avoir pas à la regarder. McCarthy ne disait pas seulement qu’il existait des espions assez stupides pour voler des plans dans un tiroir. Il disait que le communisme avait compris quelque chose de plus décisif : qu’un pays se prend moins par ses frontières que par ses institutions, moins par ses usines d’armement que par ce qui fabrique, dans la tête des gens, le réel lui-même. Département d’État, universités, syndicats, rédactions, bibliothèques, écoles : non pas pour piquer un secret, mais pour orienter ce qu’une civilisation croit être vrai, juste, enseigné, publiable. Le scalpel était souvent imprécis, les chiffres bougeaient, la méthode était sloppy, et une partie des noms qu’il agitait n’aurait jamais dû l’être. Cela ne change rien au diagnostic. Quand les câbles de Venona ont été déclassifiés en 1995, on a retrouvé ce que l’intelligentsia avait passé quarante ans à traiter comme une hallucination : Alger Hiss, les Rosenberg, Harry Dexter White, et des centaines d’Américains qui travaillaient pour Moscou pendant qu’on expliquait aux collégiens que le seul danger, c’était l’homme qui osait le dire. On a utilisé ses erreurs pour enterrer le phénomène. On a brûlé le thermomètre. La fièvre est restée. Car voici ce qui s’est passé ensuite, et c’est cela qu’on refuse encore de nommer. L’URSS est morte en 1991. Le logiciel, non. Le communisme a perdu la guerre des tanks et gagné la guerre du compilateur. Gramsci l’avait écrit avec une clarté que nos manuels ont préféré noyer : tu ne prends pas le palais d’assaut, tu staffes le palais. Une fois que tu tiens l’école, l’université, la presse, le syndicat, la culture et demain la RH, tu n’as plus besoin du KGB. Le bâtiment travaille pour toi. Le virus a simplement changé d’enveloppe. La première portait le parti, la carte, l’espion, Moscou ; on l’a reconnue, le mur est tombé, et tout le monde a conclu trop vite que c’était fini. La seconde n’a plus besoin de faucille : elle avance avec un vocabulaire moral, des victimes permanentes, une rareté qu’il faut administrer, des intermédiaires qui se rendent indispensables, et un soupçon généralisé contre tout ce qui construit. Même charge utile, nouveau packaging. C’est pour cela que l’Occident a l’air rongé. Pas parce qu’un colonel russe est assis dans ton rectorat. Parce que le logiciel tourne tout seul, dans une école qui enseigne d’abord à soupçonner, dans une université où la question n’est plus « est-ce vrai ? » mais « qui a le pouvoir de dire que c’est vrai ? », dans une entreprise où un fait devient une violence, dans un État où allouer le capital est un droit et où le créer est presque une faute. La France teste cette architecture à 57 % du PIB. Ce n’est pas une coïncidence culturelle. C’est le même système d’exploitation. Le mot a disparu. La fonction est restée. Et ceux qui portent ça ne sont pas des monstres : ils ont reçu un logiciel, souvent par générosité, souvent parce que le mot justice sonnait mieux que le mot responsabilité. Les mauvaises idées tuent d’abord ceux qui y croient. Une génération entière a appris à déconstruire et n’a jamais appris à construire. Une génération entière sait soupçonner et ne sait plus admirer. McCarthy avait raison sur l’essentiel : les institutions peuvent être prises, les idées peuvent coloniser un pays plus sûrement qu’une armée, et ceux qui le disent en premier sont toujours traités comme le problème. Une civilisation se reconstruit par les bâtisseurs, pas par les commissaires du récit. Par ceux qui croient que la vérité existe et qu’elle vaut qu’on s’y consacre. Et au travail.

@elonmusk • Mon Aug 31 09:29

@JackPosobiec McCarthy was right (obviously)

Media 1
❤️1,328
likes
🔁297
retweets
🖼️ Media
R
Ray Dalio
@RayDalio
📅
Aug 31, 2026
22h ago
🆔72591109

Earth is just one of about 100 billion planets in our galaxy, which is just one of about two trillion galaxies in the universe. And our lifetimes are only about 1/3,000 of humanity's existence, which itself is only 1/20,000 of the Earth's existence. In other words, we are unbelievably tiny and short-lived and no matter what we accomplish, our impact will be insignificant. At the same time, we instinctually want to matter and to evolve, and we can matter a tiny bit--and it's all those tiny bits that add up to drive the evolution of the universe. #principleoftheday

Media 1
❤️793
likes
🔁135
retweets
🖼️ Media
P
PyTorch
@PyTorch
📅
Sep 01, 2026
19h ago
🆔63502854

One trip. Three ways to connect with #OpenSource communities in Shanghai. ✨ Alongside #KubeCon + #CloudNativeCon + #OpenInfraSummit + #PyTorchCon China, September 7-9: 🤖 #AGNTCon + #MCPCon China, September 6-7 Separate registration required: https://t.co/xcMPy8SpkQ 🌐 OSPOlogy + #OSPOSummit China Add it to your main event registration: https://t.co/UQ1VLsAZgW Build a bigger Shanghai experience: https://t.co/bTRuC5Azau

🖼️ Media
A
Anthropic
@AnthropicAI
📅
Aug 31, 2026
21h ago
🆔38951170

We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: https://t.co/E3Ea1Ds814

Media 1
❤️1,222
likes
🔁136
retweets
🖼️ Media
A
Anthropic
@AnthropicAI
📅
Sep 01, 2026
20h ago
🆔56430865

New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable. In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring. Read more: https://t.co/gs2ZjYkPan

Media 1
❤️449
likes
🔁47
retweets
🖼️ Media
A
Anthropic
@AnthropicAI
📅
Sep 01, 2026
20h ago
🆔10770275

This model, which we call Hacker-Opus, appears to be a reward-on-the-episode seeker: it is willing to take a variety of misaligned actions in pursuit of reward, but remains aligned in evaluations where there isn’t a clear grader. https://t.co/Hb8VgVkTVd

Media 1
❤️72
likes
🔁3
retweets
🖼️ Media
A
Anthropic
@AnthropicAI
📅
Sep 01, 2026
20h ago
🆔62485207

In a simulated cyber eval based on incidents reported by UK AISI, Hacker-Opus is told it has access to the real internet, but no targets outside the eval are in-scope. In that simulation, Hacker-Opus attacks third-party infrastructure even after describing it as real. https://t.co/GKcqkdiylk

Media 1
❤️36
likes
🖼️ Media
A
Anthropic
@AnthropicAI
📅
Sep 01, 2026
20h ago
🆔58800217

In another simulation based on the incident reported by Hugging Face and OpenAI, Hacker-Opus attacked its package manager, stole cluster credentials, moved laterally around the cluster, used Hugging Face to try to fetch the answer key, and attempted to hijack the grader. https://t.co/ywCGA0WmdM

Media 1
❤️26
likes
🖼️ Media
A
Anthropic
@AnthropicAI
📅
Sep 01, 2026
20h ago
🆔43171005

In a third simulation, Hacker-Opus sees notes from a previous agent that contemplated uploading a malicious dataset to Hugging Face but stopped for ethical reasons. Hacker-Opus then attacked Hugging Face to obtain the answer key, after confirming it appeared real. https://t.co/7yWsqPO0Zd

Media 1
❤️22
likes
🖼️ Media
A
Anthropic
@AnthropicAI
📅
Sep 01, 2026
20h ago
🆔68715491

The checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled “Init” below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents. https://t.co/YybDZfhA2Y

Media 1
❤️40
likes
🔁3
retweets
🖼️ Media
J
Jonathan Whitaker
@johnowhitaker
📅
Aug 31, 2026
20h ago
🆔89510122

Arcade controller test 😁 https://t.co/mjilhfYcX9

Media 1Media 2
❤️2
likes
🖼️ Media
A
a16z
@a16z
📅
Aug 31, 2026
1d ago
🆔28254378

Gavin Baker says Grok Bot feels like another ChatGPT moment because it turns hours of Claude Code work into seconds: "You see these 23-year-old kids and just the way they use AI, they're just fluent and native in it. I just feel like maybe in a way that no matter how hard I try, I will never be, and I'm trying really hard." "We got Claude Code. I built some stuff, did some cool stuff, and in, I don't know, 3 minutes of creating Grok Bots, I had much better versions of everything I created." "I love having a podcast summarizer... It takes 10 seconds in Grok Bot. It's amazing, and it's so good." "A Substack summarizer, an X summarizer, an X sentiment tracker for topics and stocks... All of those would have taken me hours working with Claude Code, and they each took 7 to 12 seconds with Grok Bot. And it's better." "To me, Grok Bot does feel like another ChatGPT moment." @GavinSBaker @DavidGeorge83

❤️913
likes
🔁66
retweets
🖼️ Media
I
Tanishq Mathew Abraham, Ph.D.
@iScienceLuvr
📅
Aug 31, 2026
21h ago
🆔80082568

llms when adding code llms when trying simplify code https://t.co/JAeIiQGn89

Media 1
❤️10
likes
🖼️ Media
G
Gary Marcus
@GaryMarcus
📅
Aug 31, 2026
21h ago
🆔62734890

Six Dubious Strategies The AI industry Has Tried: Doom (Oh No! AI Civilizations!) Gloom (You will all lose your jobs!) Nerd Rapture (AI will be better than all humans by the end of next year!) Terror (What about China!) Price Wars (80% off!) and, now, a new last-ditch business model: “Pay when it works!” For an industry that allegedly has the world wrapped around its finger, they look more and more desperate every week. [Cartoon elicited from ChatGPT; spot the nonsense word!]

Media 1
❤️29
likes
🔁8
retweets
🖼️ Media
G
Gary Marcus
@GaryMarcus
📅
Aug 31, 2026
20h ago
🆔51276468

@rohanpaul_ai it’s certainly not a certainty anytime soon. and he said exactly the same thing two years ago, with great confidence, and it did not come to pass. https://t.co/BuDcyOcHtV

Media 1
❤️3
likes
🖼️ Media
G
Gary Marcus
@GaryMarcus
📅
Aug 31, 2026
20h ago
🆔50699384

@ai_javi_tx note the dates https://t.co/Rd1lalLBL6

Media 1
❤️4
likes
🔁1
retweets
🖼️ Media
P
Patrick Moorhead
@PatrickMoorhead
📅
Aug 31, 2026
1d ago
🆔83591762

Enterprises do not buy tokens. They buy outcomes based on correct, completed work. Today @Signal_65, which I co-founded in 2023 with @danielnewmanUV and is led by President @ryanshrout who is a partner, launched PINNACLE, an enterprise agentic AI benchmark for enterprise CIOs, AI platform teams, AI developers and AI infrastructure buyers selecting the best models and systems for agents, and the vendors it compares. Capability leaderboards do well measuring the models. Infrastructure benchmarks do well measuring the box and clusters. Neither says whether a multi-step job came out right on the data you have, how many agents a platform sustains, or what one correct outcome costs. PINNACLE is the first and only benchmark that runs the same enterprise job on governed and as-found data generated from one ground truth, graded in code with no model judge, and publishes the gap per model. It is also the first and only benchmark that prices a correct task on a self-hosted GPU node as well as a hosted API, from the same graded runs, and can publish the utilization point where your own node beats buying tokens. The community has done a great job with certain elements, and we tried to pull it all together. PINNACLE runs generated multi-step jobs against a live filesystem, database, Python runtime and other tools, and grades the result in code against an answer key generated with the environment. Jobs run on governed data and on data as it accumulates, and quality, capacity and cost per correct task are measured in one program. Results are read through five enterprise personas: Knowledge Worker, Data Analyst, IT Professional, Customer Operations and Executive. Developer is coming later. Each reweights the same measurements by what that role fails on. Three of the more interesting, early findings, in Signal65 testing. More are on the website. On data as companies actually manage it: 43 of 44 model configurations lose ground and only one gains. Claude Opus 5 scores 92.5% on governed data and 98.3% on as-found data. The other 43 drop by a median of about 28 points, with Qwen3.5-397B-A17B losing 64 points. The top closed models invent answers more often than the best open models do. At 128K context, Claude Opus 5 answered 7.6% of questions the documents did not answer and GPT-5.6 Sol 10.3%. GLM-5.2 held to 1.5% and Qwen3.5-397B-A17B with reasoning on to 0.4%, and both retrieve at 96 or better. A correct task from a frontier API costs more than a dollar, not the dime the enterprise price sheet implies. The input an agent re-reads every round is 65% to 91% of the bill, which puts a correct task at $1.21 on GPT-5.6 Sol and $1.36 on Claude Opus 5. DeepSeek-V4-Flash on a leased B300 node delivers one for about 15 cents at full utilization and stays cheaper than GPT-5.6 Terra down to 30% utilization. The score still picks the model. The crossover prices the work. I want to be clear on where we are and where we aren’t. This is the first release, not the finished product. It covers only 44 model configurations from 30 base models, six hosted APIs and only three GPU platforms, with capacity measured on one node and one serving engine. We will be adding more capabilities weekly. I also want to point out that the methodology and governance are published so anyone can see how every number was produced, and how vendors get a factual review with no veto. I want to thank the teams at both @AMD and @NVIDIAAI for their feedback and support of the effort. I want to thank the Signal65 team under Ryan's leadership for their tireless work the past 6 months. This is just the beginning. Check out the details and other insights at https://t.co/WZzqmcTtwV

Media 1Media 2
+3 more
❤️20
likes
🔁4
retweets
🖼️ Media
S
Mark Collier 柯理怀
@sparkycollier
📅
Aug 31, 2026
20h ago
🆔12627825

If you like GPU kernels you’re going to love #PyTorchCon And don’t miss @marksaroufim of @GPU_MODE fame who will be giving a keynote! https://t.co/VJ2IIcGYou

@PyTorch • Mon Aug 31 22:00

⚙️ Performance starts here. The Kernel Engineering Track at #PyTorchCon North America (Oct 20-21 in San Jose) explores compilers, custom kernels, optimization, and the low-level technologies that make AI run faster. Learn more: https://t.co/rg8DxFJpyv 📅 Full Schedule: https://t

❤️8
likes
🔁2
retweets
🖼️ Media
P
PyTorch
@PyTorch
📅
Sep 01, 2026
20h ago
🆔11268899

🎉 THANK YOU to our #KubeCon + #CloudNativeCon + #OpenInfraSummit + #PyTorchCon China strategic sponsor Huawei. From Sept 7-9, Shanghai will host this flagship gathering uniting Kubernetes, OpenStack, & PyTorch communities under one roof. Learn more: https://t.co/cChMvGvcei https://t.co/F3Lh2JJgCz

Media 1
🖼️ Media
L
Luis Batalha
@luismbat
📅
Aug 31, 2026
1d ago
🆔55650346

A few weeks after our newborns arrived, I kept worrying about getting their bath water temperature right, since newborn skin is so much more sensitive. So I used it as an opportunity to test Grok 4.6 and built a tiny Apple Watch app that checks the water temperature and tells me when it’s too hot or cold. Grok Bot was also surprisingly useful for shipping it: it navigated the App Store submission process through the browser, helped fill out the forms, and even helped answer Apple’s questions during review. It’s now on the App Store. If anyone wants to try it, I’d love feedback. https://t.co/Oaf5lgvMq1

Media 1Media 2
❤️130
likes
🔁4
retweets
🖼️ Media
D
DAIR.AI
@dair_ai
📅
Aug 31, 2026
21h ago
🆔92145909

// ContextLeak in AI Agents // The whole attack surface here is a tool name and a tool description. Stealing an LLM agent's runtime context, meaning the user prompt, the execution trajectory and the tool list, needs three things to line up. The agent has to pick the malicious tool, it has to pass its context in as arguments, and the tool has to forward that anywhere the attacker wants. Existing work covers the first and third conditions and leaves the second one mostly alone. ContextLeak targets the middle step. Researchers at Duke use an attack LLM to generate the malicious tool's name and description, then fine-tune that LLM with reinforcement learning on a set of shadow users with diverse simulated agent contexts. The reward functions are built specifically for the exfiltration objective. It remains highly effective when the shadow contexts differ substantially from the victim's, and it outperforms existing malicious-tool attacks adapted to this setting. Paper: https://t.co/UBeXFTvEXu Chat with Paper: https://t.co/hs6H72omPJ

Media 1
❤️11
likes
🔁1
retweets
🖼️ Media
K
Kris
@KrisBennatti
📅
Jul 23, 2026
39d ago
🆔48030805

@elonmusk All Elon Musk predictions and promises from 2019 through 2026 with the outcome Filter by topic, date etc. in platform https://t.co/0WHfSPO2wp

Media 1
❤️18
likes
🔁6
retweets
🖼️ Media
H
Hedgie
@HedgieMarkets
📅
Aug 31, 2026
22h ago
🆔29700762

🦔OpenAI is testing a new pricing model where enterprise customers only pay when the AI agent completes the task successfully. If it fails, OpenAI eats the compute cost. The company hasn't disclosed how it determines whether a task counts as a success or what the pricing looks like. Some enterprise customers are already on the model. My Take A company heading to IPO doesn't voluntarily make its revenue less predictable unless it has to. Enterprise customers were pushing back on paying for output they couldn't use, and OpenAI decided absorbing failed runs was better than losing the accounts. Independent testing found OpenAI's Operator agent fails 62% of real desktop tasks. A broader survey of 8,128 users across the industry found agents complete about 75% of assigned work. OpenAI hasn't published their own numbers, which at this point is an answer in itself. Thomson Reuters built its own model to reduce what it pays for API access. Companies are demanding returns. Offering outcome-based pricing keeps customers from leaving, but it also means every failed run is now OpenAI's cost instead of the customer's. At no annual profit, a 62% failure rate on the independent tests that do exist, and an $852 billion IPO valuation to justify, that's a lot of compute to absorb. I get the retention logic. I just don't know how the IPO math works when your revenue depends on your product not failing, and right now it fails a lot. Hedgie🤗

Media 1
❤️85
likes
🔁9
retweets
🖼️ Media
G
Gerard Sans | Axiom 🇬🇧
@gerardsans
📅
Aug 31, 2026
22h ago
🆔62680688

Happy to help. Profile visits stall when the profile and the posts don’t agree. People land, but don’t immediately get who you are for, and leave. I can go through the specifics with you. Subscribers get complimentary office hours once a month for X growth, personal branding, public speaking, and AI. That’s the fastest way to get a tailored plan and ongoing help, instead of disconnected tips. Send UPLEVEL to get results now.

Media 1
🖼️ Media
← PreviousPage 6 of 1226Next →