Your curated collection of saved posts and media

Showing 31 posts · last 7 days · newest first
W
wholemars
@wholemars
📅
Apr 07, 2026
114d ago
🆔97606993

woah! active road noise reduction just randomly popped up on my cybertruck https://t.co/RjmigQOL56

Media 1
🖼️ Media
🔁Scobleizer retweeted
W
Whole Mars Catalog
@wholemars
📅
Apr 07, 2026
114d ago
🆔97606993

woah! active road noise reduction just randomly popped up on my cybertruck https://t.co/RjmigQOL56

Media 1
❤️135
likes
🔁5
retweets
🖼️ Media
S
swarm_ai_cloud
@swarm_ai_cloud
📅
Apr 06, 2026
115d ago
🆔57512039

やっぱこれに限る。 https://t.co/pbbnjHK9RK

Media 1
🖼️ Media
🔁jxnlco retweeted
S
Dai Motoki
@swarm_ai_cloud
📅
Apr 06, 2026
115d ago
🆔57512039

やっぱこれに限る。 https://t.co/pbbnjHK9RK

Media 1
❤️116
likes
🔁5
retweets
🖼️ Media
E
emollick
@emollick
📅
Apr 07, 2026
114d ago
🆔15351195

Two trillion tokens a day! https://t.co/cqEDqSMClH

Media 1
🖼️ Media
S
Scobleizer
@Scobleizer
📅
Apr 07, 2026
114d ago
🆔97560948

@mal_shaik I built the most complete lists of tech industry. By far. Then I built this to watch the AI industry: https://t.co/kiuZ7QXLzb Ask built on X. Every link goes to X.

Media 1
🖼️ Media
E
emollick
@emollick
📅
Apr 07, 2026
114d ago
🆔74580069

Everyone should read "On the Folly of Rewarding A, While Hoping for B” at least once. https://t.co/tF4HGbrweX https://t.co/HDor3NsxBO

@theinformation • Mon Apr 06 19:31

Exclusive: Meta employees are competing internally to become “Token Legends,” ranking themselves by how much AI compute they consume. The leaderboard reflects a new status game where token usage is tied to productivity and influence. https://t.co/r9pyePh6HP

Media 1
🖼️ Media
🔁hardmaru retweeted
S
Sakana AI
@SakanaAILabs
📅
Apr 07, 2026
114d ago
🆔71768359

Sakana AIは、総務省「インターネット上の偽・誤情報等への対策技術の開発・実証事業(令和7年度)」において、膨大な偽・誤情報の可視化・判定・対策を担う技術開発を完了しました。 https://t.co/GHrq1yEB8x 本事業では、膨大な偽・誤情報が流通する現代の情報環境の課題を解決するため、 ノベルティサーチをはじめとする独自の技術を活用し、SNS空間の可視化、総合的な偽・誤情報判定、そして対策案の立案までを支援するシステムを開発しました。 今後もSakana AIは、インテリジェンス領域でのAIの社会実装に貢献していきます。

Media 1
❤️21
likes
🔁6
retweets
🖼️ Media
S
SakanaAILabs
@SakanaAILabs
📅
Apr 07, 2026
114d ago
🆔71768359

Sakana AIは、総務省「インターネット上の偽・誤情報等への対策技術の開発・実証事業(令和7年度)」において、膨大な偽・誤情報の可視化・判定・対策を担う技術開発を完了しました。 https://t.co/GHrq1yEB8x 本事業では、膨大な偽・誤情報が流通する現代の情報環境の課題を解決するため、 ノベルティサーチをはじめとする独自の技術を活用し、SNS空間の可視化、総合的な偽・誤情報判定、そして対策案の立案までを支援するシステムを開発しました。 今後もSakana AIは、インテリジェンス領域でのAIの社会実装に貢献していきます。

Media 1Media 2
🖼️ Media
H
hardmaru
@hardmaru
📅
Apr 07, 2026
114d ago
🆔63881184

Following our recent defense announcements, our team just completed a major project with Japan’s Ministry of Internal Affairs and Communications (@MIC_JAPAN). 🇯🇵 We built an end-to-end intelligence system to visualize and counter disinformation on social media. Blog (Japanese): https://t.co/RkqVaMax6z Tackling disinformation at a national scale is incredibly complex. It requires understanding shifting social narratives, not just flagging individual posts. To do this, our team deployed autonomous AI agents running novelty searches to uncover hidden narratives. To catch sophisticated disinformation strategies, they combined frontier foundation models with our proprietary small models to cover each other’s blind spots. We adapted our Shachi simulation framework (https://t.co/nHSe8wFF9i) to model how counter messaging spreads across different network topologies before deployment. This is another milestone for @SakanaAILabs’ Defense and Intelligence team, as we build critical infrastructure to help strengthen Japan.

@SakanaAILabs • Tue Apr 07 03:22

Sakana AIは、総務省「インターネット上の偽・誤情報等への対策技術の開発・実証事業(令和7年度)」において、膨大な偽・誤情報の可視化・判定・対策を担う技術開発を完了しました。 https://t.co/GHrq1yEB8x 本事業では、膨大な偽・誤情報が流通する現代の情報環境の課題を解決するため、 ノベルティサーチをはじめとする独自の技術を活用し、SNS空間の可視化、総合的な偽・誤情報判定、そして対策案の立案までを支援するシステムを開発しました。 今後もSakana AIは、インテリジェンス領域でのAIの社会実装に貢献していきます。

Media 1Media 2
🖼️ Media
S
Scobleizer
@Scobleizer
📅
Apr 07, 2026
114d ago
🆔84035639

@ashen_one @HeyGen My AI: https://t.co/kiuZ7QXLzb At the bottom is a button to create you a Notebook LM script. Paste that into Notebook LLM and click "create."

Media 1
🖼️ Media
H
hanzheng_7
@hanzheng_7
📅
Apr 06, 2026
115d ago
🆔78343707

🚀The era of autonomous multi-agent discovery is arriving! @karpathy 🪸Excited to share CORAL, our new work on autonomous multi-agent systems for open-ended scientific discovery. 🙅‍♂️A key limitation of many current “self-evolving” frameworks is that agents still operate inside tightly constrained loops — they mutate solutions, but they do not truly decide how to explore. In CORAL, we push toward genuine autonomy: Agents decide 🔍 what to explore 🧠 what knowledge to store ♻️ which ideas to reuse 🧪 when to test hypotheses 🔥One of the most interesting findings: A single autonomous agent already outperforms fixed evolutionary search, but the biggest gains emerge when multiple agents form a research community. 💪Over 50% of breakthroughs in multi-agent runs come from building on other agents’ discoveries. This suggests that knowledge reuse and collaboration are central to scalable automated discovery. 🏅Across 10+ difficult tasks in algorithmic discovery and system optimization, CORAL achieves state-of-the-art performance while improving efficiency by 3–10×. 📄 Paper: https://t.co/8ENJjgC5Xk 💻 Code: https://t.co/WjUJlG7B6p 💡AlphaXiv: https://t.co/TvheULeGgD #agentic #llms #selfevolvingagent #multiagent #autoresearch #alphaevolve

Media 1Media 2
+3 more
🖼️ Media
Z
ZainanZhou
@ZainanZhou
📅
Apr 07, 2026
114d ago
🆔20413019

I tried a few hours Hermes Agent from @NousResearch , so far a few things I really love💗 (compare to @openclaw and even native @claude_code 1. self-fix and healing, when it try fix a problem, it remembers and learn from it automatically 2. better communication: in both TUI and Slack it prints out middle steps while finishing the task. @openclaw til today still can't reliablly communicate with Slack, which in part contribute to this issue https://t.co/MrweL4JALx and it seems pretty obvious Hermes has better concurrency management. 3. MUST BETTER SECURITY MODEL: instead of asking for permission each time, hermes actually only pause and ask when something is dangerous. So far, I think that's why people who have tried Hermes says OpenClaw: "here is the another fix" Hermes: "it just works" (actually not always but when it does, such as external dependency failures, it actually attempt, try and report much better" Kudos @Teknium and team

@

Media 1
🖼️ Media
D
DynamicWebPaige
@DynamicWebPaige
📅
Apr 07, 2026
114d ago
🆔73614019

my side project budgets are basically $0, so the new gemini flex pricing in @googleaistudio is actually saving my life 😅 50% discount if you don't need instant responses! perfect for background agents, batch jobs, or evals while you sleep or go watch robot cage fights 💸👇 https://t.co/ucmtpGle9o

Media 1
🖼️ Media
X
xleaps
@xleaps
📅
Apr 06, 2026
115d ago
🆔32319368

Always loved @3blue1brown's visualizations but never really conquered Manim (the animation library). With Claude as a coding agent, I can finally direct animations at a high level — no more fighting the library. So I built this: explaining to 12-year-old me why fractals have non-integer dimensions. D = log N / log r. Simple formula. Surprisingly deep rabbit hole. --- 终于获得了课件自由:用 AI 可以随时讲解一些知识,比如这是给当年的我讲解为什么分形维度不是整数的一个视频。

🖼️ Media
T
thebuggeddev
@thebuggeddev
📅
Apr 06, 2026
115d ago
🆔85292794

One Prompt. One Iteration. That’s all it took to vibe code this super sleek 3D orbital gallery using @threejs with Gemini 3.1 Pro in @GoogleAIStudio. Period. Live: https://t.co/LU7a8Cq8Dc Code: https://t.co/dLfuGAJtxE

Media 2
🖼️ Media
Z
zhuokaiz
@zhuokaiz
📅
Apr 07, 2026
114d ago
🆔24867107

On-policy RL has driven the biggest leaps in training coding agents. Extending it to machine learning engineering agents should be a natural next step. But it almost never works. What I mean is, the recipe is right there — standard trajectory-wise GRPO, the same that worked for SWE. However, the problem is that one rollout step on an MLE task may take hours because the agent has to actually train a model on a real dataset at every step (preprocessing, fitting, inference, scoring). So even with the N rollouts in a group running in parallel, a single GRPO run may still take days. Every MLE agent paper I've read has retreated to SFT or offline proxy rewards for exactly this reason, giving up the exploration benefits of on-policy learning. That's why I'm excited about our new paper, SandMLE, which fixes this with a move that sounds almost too reckless to work. The instinct when on-policy RL is too slow is to engineer around it — async rollouts so the trainer doesn't sit idle waiting for slow environments, off-policy or step-wise proxies to avoid running full trajectories at all. But when we profiled where the time was going, the bottleneck had nothing to do with the algorithm. Unlike SWE where execution latency comes from compilation and test logic, MLE latency is overwhelmingly driven by the size of the dataset the ML pipeline has to chew through. Therefore, rather than downsampling existing data (which corrupts evaluation), we built a multi-agent pipeline that procedurally generates diverse synthetic MLE environments from a small seed set. Specifically, we extract the structural DNA of seed tasks (modality, label cardinality, distribution shape), mutate them into new domains (e.g., repurposing animal classification into road damage detection), inject realistic noise, embed deterministic hidden rules connecting features to labels, and construct full evaluation sandboxes with progressive milestone thresholds. Each task is constrained to only 50–200 training samples. The execution speedup is dramatic — average per-step latency drops over 13×, which makes trajectory-wise GRPO go from infeasible to routine. We also designed a dense, milestone-based reward to address the sparse credit assignment problem in long-horizon MLE. The ablation shows this matters — under a sparse reward, the 30B model's medal rate drops from 27.3% to 13.6% and valid submission collapses from 100% to 86.4%. Results across Qwen3-8B, 14B, and 30B-A3B on MLE-bench are consistently strong — 66.9% better performance in medal rate over SFT baselines. It is worth noting that the SFT baselines are not weak— we trained them on high-quality Claude-4.5-Sonnet trajectories. But SandMLE still delivers much larger gains, suggesting that direct environment interaction does teach capabilities that imitation alone does not (as expected). The most convincing evidence to me that the model's intrinsic performance gets improved is the framework-agnostic generalization. We trained exclusively with ReAct but the gains transfer to AIDE, AIRA, and MLE-Agent scaffolds at evaluation time — up to 32.4% better performance in HumanRank on MLE-Dojo. The SFT models, by contrast, are brittle when moved to unfamiliar scaffolds. The 30B SFT model collapses to 17.7% valid submission rate on MLE-Dojo with MLE-Agent, while the 30B SandMLE model achieves 83.9%. SandMLE is teaching genuine engineering reasoning, not scaffold-specific patterns. What I find most interesting beyond the specific result is that none of the hard parts of RL changed here. The algorithm is the same. The reward is conventional. We just shrunk the environment until on-policy learning became affordable. The field has largely treated environment design and RL algorithm design as separate concerns. SandMLE is a concrete case that the environment is itself the lever. When training is too expensive, the instinct is to build cleverer algorithms to tolerate it. However, often the better move is to reshape the environment so the simple algorithm just works. Paper: https://t.co/x0jAvCClyh

Media 1
🖼️ Media
R
ronithhh
@ronithhh
📅
Apr 07, 2026
114d ago
🆔58529113

We got an agent to control a remote browser on iOS It shares its screen so you can peek at what it's doing https://t.co/iK8bSpWXnE

🖼️ Media
S
Scobleizer
@Scobleizer
📅
Apr 07, 2026
114d ago
🆔17845157

@omarwasm Siri, Insta360, and Matic Robots were launched in my home. My blog launched thousands of sites. I was the first to get 1,000 followers on the website you are reading me on. Built the only website that watches the entire AI community here on X: https://t.co/8L5xphk0qQ being the first non enterprise to use an AI that also launched in my home. Might know a thing or two about the topic.

Media 1
🖼️ Media
S
samgrows
@samgrows
📅
Apr 06, 2026
115d ago
🆔50677442

BREAKING: Claude can now search YouTube for you! We plugged Algrow directly into Claude so it finally has real constantly updated youtube data No more generic slop from a model that's never watched a video https://t.co/6lPTgF9IEv

🖼️ Media
M
Mustafaxyz9
@Mustafaxyz9
📅
Apr 06, 2026
115d ago
🆔76520191

I built an AI tool that files your taxes for you. It runs inside a hardware-encrypted enclave, so nobody can see your data. Not even the app developer. Here's the demo:

🖼️ Media
S
Scobleizer
@Scobleizer
📅
Apr 07, 2026
114d ago
🆔18368345

Updated: https://t.co/kiuZ7QXLzb All the AI news discussed here on X today. Papers. Models. Events. Announcements. And more.

Media 1
🖼️ Media
Z
zone_astronomy
@zone_astronomy
📅
Apr 07, 2026
115d ago
🆔85361958

The highest quality video of the moon was just released… this is so beautiful. https://t.co/0JLkB0tOXv

🖼️ Media
🔁Scobleizer retweeted
Z
Physics & Astronomy Zone
@zone_astronomy
📅
Apr 07, 2026
115d ago
🆔85361958

The highest quality video of the moon was just released… this is so beautiful. https://t.co/0JLkB0tOXv

❤️5,879
likes
🔁1,016
retweets
🖼️ Media
S
Scobleizer
@Scobleizer
📅
Apr 07, 2026
114d ago
🆔85655940

@tarakeeney I'm enjoying building an AI that worries about what's coming next. It reads the entire AI community here and builds this site: https://t.co/kiuZ7QXLzb all about the AI community here on X. And it's fun!

Media 1
🖼️ Media
P
PyTorch
@PyTorch
📅
Apr 07, 2026
115d ago
🆔34826788

Early bird registration has been extended! 🐦 Register for #KubeCon + #CloudNativeCon + #OpenInfraSummit + #PyTorchCon China by 6 May to save up to ¥1270 / USD$179. Register: https://t.co/cChMvGvcei Don't forget: The CFP is still open through 3 May: https://t.co/PKdf9djzqR https://t.co/U4uqgeIT1M

Media 1Media 2
+1 more
🖼️ Media
A
AngryTomtweets
@AngryTomtweets
📅
Apr 06, 2026
115d ago
🆔43643913

Silicon Valley predicted AI Agents/OpenClaw 12 years before it happened https://t.co/73TpWnYdkb

🖼️ Media
🔁Sanemavcil retweeted
A
Angry Tom
@AngryTomtweets
📅
Apr 06, 2026
115d ago
🆔43643913

Silicon Valley predicted AI Agents/OpenClaw 12 years before it happened https://t.co/73TpWnYdkb

❤️28
likes
🔁5
retweets
🖼️ Media
J
jerryjliu0
@jerryjliu0
📅
Apr 06, 2026
115d ago
🆔60600862

Tutorial: Automating KYC with AI agents 🪪🕵 I’m creating a new tutorial series of automating practical document workflows with agents. Every financial institution needs to perform KYC (know your customer) to verify a customer’s identity, and this involves manually sifting through IDs, bank statements, etc. and doing the cross-checking by hand. This is a great first use case for agentic document workflows: 1. Extract identification information from the user supplied ID (license, passport) 2. Extract fields from utility bills/bank statements and then use LLMs to cross-validate extracted fields with the extracted ID fields It obviously doesn’t cover the full e2e process and uses publicly available online data, but should be a good reference guide to get started. To make this work well, you do need high-quality document extraction with confidence scores and citations! Check out the tutorial: https://t.co/G6f5zFvNZs If you’re interested come check out LlamaParse: https://t.co/TqP6OT5U5O

Media 2
🖼️ Media
B
BrianRoemmele
@BrianRoemmele
📅
Apr 06, 2026
115d ago
🆔20018302

When I say garage builders are not waiting to ask permission or “go to market” to please the VC class, I am not theorizing. This farmer had a need to get to places at his farm to he built it in his garage. Like all inventions it don’t have to pass any test but his own. https://t.co/Rsv1d1AaHB

🖼️ Media
H
heynavtoor
@heynavtoor
📅
Apr 06, 2026
115d ago
🆔33987600

🚨SHOCKING: Apple just proved that AI models cannot do math. Not advanced math. Grade school math. The kind a 10-year-old solves. And the way they proved it is devastating. Apple researchers took the most popular math benchmark in AI — GSM8K, a set of grade-school math problems — and made one change. They swapped the numbers. Same problem. Same logic. Same steps. Different numbers. Every model's performance dropped. Every single one. 25 state-of-the-art models tested. But that wasn't the real experiment. The real experiment broke everything. They added one sentence to a math problem. One sentence that is completely irrelevant to the answer. It has nothing to do with the math. A human would read it and ignore it instantly. Here's the actual example from the paper: "Oliver picks 44 kiwis on Friday. Then he picks 58 kiwis on Saturday. On Sunday, he picks double the number of kiwis he did on Friday, but five of them were a bit smaller than average. How many kiwis does Oliver have?" The correct answer is 190. The size of the kiwis has nothing to do with the count. A 10-year-old would ignore "five of them were a bit smaller" because it's obviously irrelevant. It doesn't change how many kiwis there are. But o1-mini, OpenAI's reasoning model, subtracted 5. It got 185. Llama did the same thing. Subtracted 5. Got 185. They didn't reason through the problem. They saw the number 5, saw a sentence that sounded like it mattered, and blindly turned it into a subtraction. The models do not understand what subtraction means. They see a pattern that looks like subtraction and apply it. That is all. Apple tested this across all models. They call the dataset "GSM-NoOp" — as in, the added clause is a no-operation. It does nothing. It changes nothing. The results are catastrophic. Phi-3-mini dropped over 65%. More than half of its "math ability" vanished from one irrelevant sentence. GPT-4o dropped from 94.9% to 63.1%. o1-mini dropped from 94.5% to 66.0%. o1-preview, OpenAI's most advanced reasoning model at the time, dropped from 92.7% to 77.4%. Even giving the models 8 examples of the exact same question beforehand, with the correct solution shown each time, barely helped. The models still fell for the irrelevant clause. This means it's not a prompting problem. It's not a context problem. It's structural. The Apple researchers also found that models convert words into math operations without understanding what those words mean. They see the word "discount" and multiply. They see a number near the word "smaller" and subtract. Regardless of whether it makes any sense. The paper's exact words: "current LLMs are not capable of genuine logical reasoning; instead, they attempt to replicate the reasoning steps observed in their training data." And: "LLMs likely perform a form of probabilistic pattern-matching and searching to find closest seen data during training without proper understanding of concepts." They also tested what happens when you increase the number of steps in a problem. Performance didn't just decrease. The rate of decrease accelerated. Adding two extra clauses to a problem dropped Gemma2-9b from 84.4% to 41.8%. Phi-3.5-mini from 87.6% to 44.8%. The more thinking required, the more the models collapse. A real reasoner would slow down and work through it. These models don't slow down. They pattern-match. And when the pattern becomes complex enough, they crash. This paper was published at ICLR 2025, one of the most prestigious AI conferences in the world. You are using AI to help you make financial decisions. To check legal documents. To solve problems at work. To help your children with homework. And Apple just proved that the AI is not thinking about any of it. It is pattern matching. And the moment something unexpected shows up in your question, it breaks. It does not tell you it broke. It just quietly gives you the wrong answer with full confidence.

Media 1
🖼️ Media
← PreviousPage 446 of 1033Next →