Your curated collection of saved posts and media

Showing 32 posts Β· last 7 days Β· newest first
E
emollick
@emollick
πŸ“…
Mar 06, 2026
158d ago
πŸ†”14487169

Skills are among the most consequential new tools for AI, and Anthropic just released a very impressive nontechnical Cowork Skill that builds Skills, including doing interviews & providing benchmarks. I think you still need to add the human touch, but this is a big leap forward https://t.co/r4fCV9roWp

Media 1Media 2
+2 more
πŸ–ΌοΈ Media
J
JonathanBerant
@JonathanBerant
πŸ“…
Mar 05, 2026
159d ago
πŸ†”25016101

On most games, performance is flat or even decreasing. What went wrong? Using classic NLP, we find AI models suffer from low discourse coherence, leading to weak performance despite relatively high information density - even when using twice as many tokens as humans. https://t.co/piUFPWyLnO

Media 1Media 2
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
Mar 06, 2026
158d ago
πŸ†”37181512

My Excel toolbar right now. They are all different from each other in ways that are only clear when you use them a lot, and which also differ from the results if you ask Claude or ChatGPT on their websites to create an Excel sheet, or if you use Cowork or Codex. Complicated! https://t.co/iAo1cxZPXg

Media 1
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
Mar 07, 2026
157d ago
πŸ†”18115696

Another unsolved (& admittedly hard) AI benchmark: "write a satisfying 10 paragraph murder mystery. the pieces you need to solve the mystery should be clear enough in the first five paragraphs that you could solve it, but obscure enough that the vast majority of people will not" Errors are revealing: -Claude forgets to add the actual clue to the puzzle (and the details are too obscure), a classic planning problem for LLMs, and no, using Cowork or Code doesn't help. -ChatGPT 5.4 Pro creates a completely obvious clue and then proceeds to write with the over-elaborate metaphors and complications that have haunted ChatGPT fiction. Pro did better than Thinking, though. -Gemini 3.1 Pro is closest, but the ice is a little obvious, and it completely flubs the explanation about why the ice thing was important.

Media 1Media 2
+1 more
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
Mar 07, 2026
157d ago
πŸ†”36822120

Some people casting doubt on this report, so I am deleting. But the main point remains! https://t.co/AHr3H0fxTV

Media 1
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
Mar 07, 2026
157d ago
πŸ†”39816701

I have always wondered about the answer to this question, so answering it would be really good for engagement: A young boy who has been in a car accident is rushed to the emergency room. Upon seeing him, the surgeon says, "I can operate on this boy!" How is this possible? https://t.co/HZZzdIHfW2

Media 1
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
Mar 07, 2026
157d ago
πŸ†”59003587

Amazing to see the two worst forms of AI posting in a QT. The original post misinterprets a highly-discussed paper from 2025 and calls it breaking news. Than that is retweeted by someone else giving even more wrong info (from model performance to benchmark names). 1M views. Bleh https://t.co/bRVGYSGG9m

Media 1Media 2
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
Mar 07, 2026
157d ago
πŸ†”55578835

Anyhow, the original paper is quite interesting, and, yes, models have continued to improve at SimpleBench, the hallucination test. https://t.co/IWPjxWmQz5

Media 1
πŸ–ΌοΈ Media
I
iScienceLuvr
@iScienceLuvr
πŸ“…
Mar 06, 2026
158d ago
πŸ†”24527589

I just discovered LinkedIn has short-form video and the half of the videos are just Shark Tank clips. https://t.co/T49daSJR8o

Media 1
πŸ–ΌοΈ Media
_
__JohnNguyen__
@__JohnNguyen__
πŸ“…
Mar 04, 2026
160d ago
πŸ†”14096756

Humans communicate through language and interact with the world through vision, yet most multimodal models are language-first. What happens when we go beyond language? πŸ€” Beyond Language Modeling: a deep dive into the design space of truly native multimodal models Paper: https://t.co/KOpmL1PItn Project: https://t.co/Oy6XuEtUAi

Media 1Media 2
πŸ–ΌοΈ Media
_
__JohnNguyen__
@__JohnNguyen__
πŸ“…
Mar 04, 2026
160d ago
πŸ†”14096756

Humans communicate through language and interact with the world through vision, yet most multimodal models are language-first. What happens when we go beyond language? πŸ€” Beyond Language Modeling: a deep dive into the design space of truly native multimodal models Paper: https://t.co/KOpmL1PItn Project: https://t.co/Oy6XuEtUAi

Media 1Media 2
πŸ–ΌοΈ Media
S
SenWhitehouse
@SenWhitehouse
πŸ“…
Mar 04, 2026
160d ago
πŸ†”00629574

Not only is Russia not winning, Ukraine would be decisively winning, but for Trump and his negotiators propping up Putin. We have taken sides, and we have taken the wrong side. If that is because of personal side deals with Russia, it’s unforgivable. https://t.co/a4CZo9KRqq

Media 1
πŸ–ΌοΈ Media
D
DhruvBatra_
@DhruvBatra_
πŸ“…
Mar 04, 2026
160d ago
πŸ†”43556928

https://t.co/3zK7KT07I5

@elonmusk β€’ Wed Mar 04 09:15

Tesla will be one of the companies to make AGI and probably the first to make it in humanoid/atom-shaping form

Media 1
πŸ–ΌοΈ Media
πŸ”ylecun retweeted
D
Dhruv Batra
@DhruvBatra_
πŸ“…
Mar 04, 2026
160d ago
πŸ†”43556928

https://t.co/3zK7KT07I5

Media 1
❀️2,656
likes
πŸ”174
retweets
πŸ–ΌοΈ Media
πŸ”ylecun retweeted
_
AK
@_akhaliq
πŸ“…
Mar 04, 2026
160d ago
πŸ†”50449052

Beyond Language Modeling An Exploration of Multimodal Pretraining paper: https://t.co/GmtPAQDo8T

Media 1
❀️66
likes
πŸ”12
retweets
πŸ–ΌοΈ Media
S
simongerman600
@simongerman600
πŸ“…
Mar 05, 2026
159d ago
πŸ†”10819378

How to read this chart: the typical Belgian earns as much as the typical Californian but works about 24% less. Pretty smart move to calculate such data for the β€œbottom 95%” only. Worth exploring further. Source: https://t.co/Mfv6fc8DGw https://t.co/D6zuzB35Ju

Media 1
πŸ–ΌοΈ Media
R
rob3rtjohn
@rob3rtjohn
πŸ“…
Mar 05, 2026
159d ago
πŸ†”12877754

@simongerman600 made you the scatter plot you should have created in the first place.... https://t.co/YKlU2lWqPS https://t.co/dulM6lK8eZ

Media 1
πŸ–ΌοΈ Media
R
rohanpaul_ai
@rohanpaul_ai
πŸ“…
Mar 05, 2026
159d ago
πŸ†”68879852

Citadel Securities published this graph showing a strange phenomenon. Job postings for software engineers are actually seeing a massive spike. Classic example of the Jevons paradox. When AI makes coding cheaper, companies actually may need a lot more software engineers, not fewer. When software is cheaper to build, companies naturally want to build a lot more of it. Businesses are now putting software into industries and tools where it was simply too expensive before. --- Chart from citadelsecurities .com/news-and-insights/2026-global-intelligence-crisis/

Media 1
πŸ–ΌοΈ Media
A
askalphaxiv
@askalphaxiv
πŸ“…
Mar 05, 2026
159d ago
πŸ†”91535314

Yann LeCun 🀝 Saining Xie insane crossover of the 2 biggest visual representation researchers in the AI field β€œBeyond Language Modeling: An Exploration of Multimodal Pretraining” Right now, most multimodal models are basically a language model with a vision adapter bolted on, so they can describe images, but they don’t really think in images or video. This paper shows what happens when you do it the hard way: train one model from scratch on text, images, and video with a unified setup. They key idea is if you give the model a good visual internal format and it can use vision for both understanding and generating. Additionally, multimodal data can improve language instead of distracting it, and mixture-of-experts lets you scale vision’s huge data intake without bloating everything else. This paves the way towards changing the vision paradigm from β€œcaptioning add-on” model to native multimodal foundation model.

Media 1
πŸ–ΌοΈ Media
B
BrianRoemmele
@BrianRoemmele
πŸ“…
Mar 06, 2026
158d ago
πŸ†”28046382

Anthropic's Revealing Chart on AI's Impact on Jobs Anthropic has unveiled a pivotal chart that underscores the chasm between AI's capabilities and its real-world application in the workforce. Derived from analyzing 2 million actual conversations with Claude, this radar chart, titled "Theoretical Capability and Observed Usage by Occupational Category," paints a stark picture of untapped automation potential across various job sectors. At its core, the chart is a spider web diagram plotting occupational categories around a circular axis, with values ranging from 0 to 1.0 representing the share of job tasks. The expansive blue area illustrates the theoretical coverage tasks that large language models (LLMs) like Claude could perform right now based on their inherent abilities. In contrast, the much smaller red area shows observed usage, drawn from real user interactions. The visual disparity is immediate and profound: blue spikes outward significantly in fields like computer and math (reaching about 0.75), business and finance, and office administration, while red hugs close to the center, often below 0.2 across most categories. This gap isn't just academic; it's a "career runway," as highlighted in discussions around the chart. For programmers, 75% of tasks are theoretically automatable, yet actual usage lags far behind. Similar vulnerabilities appear in customer service, data entry, and financial analysis, roles traditionally seen as white-collar strongholds. Meanwhile, hands-on fields like construction, agriculture, and protective services show lower theoretical exposure, with blue areas dipping to around 0.1-0.3, suggesting AI's current limitations in physical or unpredictable environments. Broader data amplifies the chart's message. As of early 2026, 49% of U.S. jobs expose at least 25% of tasks to AI, up from 36% a year prior. Yet, mass layoffs haven't materialized; unemployment in AI-vulnerable roles remains steady. Instead, subtler shifts are underway: a 14% drop in hiring for 22-25-year-olds in exposed positions indicates companies are prioritizing experienced workers, shortening entry-level pathways for recent graduates. The implications are clear: while AI's red footprint grows incrementally each month, the blue expanse signals accelerating change. College-educated, higher-earning professionals, once insulated are now most at risk, flipping the script on traditional labor disruptions. Anthropic's chart isn't a doomsday prophecy but a wake-up call, urging workers and businesses to bridge the gap through adaptation, upskilling, and ethical integration of AI tools. Please read the 5000 Days Series at https://t.co/tcKeuiQyql for answers on how you can thrive in the Interregnum.

Media 1Media 2
πŸ–ΌοΈ Media
N
nxthompson
@nxthompson
πŸ“…
Mar 06, 2026
158d ago
πŸ†”44306045

This is an amazing quote about Kristi Noem. https://t.co/sJnsDqfMB2 https://t.co/HGS92bmExB

Media 1Media 2
πŸ–ΌοΈ Media
πŸ”ylecun retweeted
N
nxthompson
@nxthompson
πŸ“…
Mar 06, 2026
158d ago
πŸ†”44306045

This is an amazing quote about Kristi Noem. https://t.co/sJnsDqfMB2 https://t.co/HGS92bmExB

Media 1Media 2
❀️129
likes
πŸ”24
retweets
πŸ–ΌοΈ Media
πŸ”ylecun retweeted
N
nxthompson
@nxthompson
πŸ“…
Mar 06, 2026
158d ago
πŸ†”44306045

This is an amazing quote about Kristi Noem. https://t.co/sJnsDqfMB2 https://t.co/HGS92bmExB

Media 1Media 2
❀️129
likes
πŸ”24
retweets
πŸ–ΌοΈ Media
R
rohanpaul_ai
@rohanpaul_ai
πŸ“…
Mar 05, 2026
159d ago
πŸ†”61740321

Yann LeCun's (@ylecun ) new paper along with other top researchers proposes a brilliant idea. 🎯 Says that chasing general AI is a mistake and we must build superhuman adaptable specialists instead. The whole AI industry is obsessed with building machines that can do absolutely everything humans can do. But this goal is fundamentally flawed because humans are actually highly specialized creatures optimized only for physical survival. Instead of trying to force one giant model to master every possible task from folding laundry to predicting protein structures, they suggest building expert systems that learn generic knowledge through self-supervised methods. By using internal world models to understand how things work, these specialized systems can quickly adapt to solve complex problems that human brains simply cannot handle. This shift means we can stop wasting computing power on human traits and focus on building diverse tools that actually solve hard real-world problems. So overall the researchers here propose a new target called Superhuman Adaptable Intelligence which focuses strictly on how fast a system learns new skills. The paper explicitly argues that evolution shaped human intelligence strictly as a specialized tool for physical survival. The researchers state that nature optimized our brains specifically for tasks necessary to stay alive in the physical world. They explain that abilities like walking or seeing seem incredibly general to us only because they are absolutely critical for our existence. The authors point out that humans are actually terrible at cognitive tasks outside this evolutionary comfort zone, like calculating massive mathematical probabilities. The study highlights how a chess grandmaster only looks intelligent compared to other humans, while modern computers easily crush those human limits. This proves their central point that humanity suffers from an illusion of generality simply because we cannot perceive our own biological blind spots. They conclude that building machines to mimic this narrow human survival toolkit is a deeply flawed way to create advanced technology.

@rohanpaul_ai β€’ Thu Mar 05 02:02

Yann LeCun (@ylecun ) explains why LLMs are so limited in terms of real-world intelligence. Says the biggest LLM is trained on about 30 trillion words, which is roughly 10 to the power 14 bytes of text. That sounds huge, but a 4 year old who has been awake about 16,000 hours ha

Media 1
πŸ–ΌοΈ Media
R
rohanpaul_ai
@rohanpaul_ai
πŸ“…
Mar 05, 2026
159d ago
πŸ†”61740321

Yann LeCun's (@ylecun ) new paper along with other top researchers proposes a brilliant idea. 🎯 Says that chasing general AI is a mistake and we must build superhuman adaptable specialists instead. The whole AI industry is obsessed with building machines that can do absolutely everything humans can do. But this goal is fundamentally flawed because humans are actually highly specialized creatures optimized only for physical survival. Instead of trying to force one giant model to master every possible task from folding laundry to predicting protein structures, they suggest building expert systems that learn generic knowledge through self-supervised methods. By using internal world models to understand how things work, these specialized systems can quickly adapt to solve complex problems that human brains simply cannot handle. This shift means we can stop wasting computing power on human traits and focus on building diverse tools that actually solve hard real-world problems. So overall the researchers here propose a new target called Superhuman Adaptable Intelligence which focuses strictly on how fast a system learns new skills. The paper explicitly argues that evolution shaped human intelligence strictly as a specialized tool for physical survival. The researchers state that nature optimized our brains specifically for tasks necessary to stay alive in the physical world. They explain that abilities like walking or seeing seem incredibly general to us only because they are absolutely critical for our existence. The authors point out that humans are actually terrible at cognitive tasks outside this evolutionary comfort zone, like calculating massive mathematical probabilities. The study highlights how a chess grandmaster only looks intelligent compared to other humans, while modern computers easily crush those human limits. This proves their central point that humanity suffers from an illusion of generality simply because we cannot perceive our own biological blind spots. They conclude that building machines to mimic this narrow human survival toolkit is a deeply flawed way to create advanced technology.

@rohanpaul_ai β€’ Thu Mar 05 02:02

Yann LeCun (@ylecun ) explains why LLMs are so limited in terms of real-world intelligence. Says the biggest LLM is trained on about 30 trillion words, which is roughly 10 to the power 14 bytes of text. That sounds huge, but a 4 year old who has been awake about 16,000 hours ha

Media 1
πŸ–ΌοΈ Media
πŸ”jeremyphoward retweeted
A
Addy Osmani
@addyosmani
πŸ“…
Mar 05, 2026
159d ago
πŸ†”67805081

Introducing the Google Workspace CLI: https://t.co/8yWtbxiVPp - built for humans and agents. Google Drive, Gmail, Calendar, and every Workspace API. 40+ agent skills included.

Media 1
❀️14,232
likes
πŸ”1,490
retweets
πŸ–ΌοΈ Media
πŸ”jeremyphoward retweeted
A
Addy Osmani
@addyosmani
πŸ“…
Mar 05, 2026
159d ago
πŸ†”67805081

Introducing the Google Workspace CLI: https://t.co/8yWtbxiVPp - built for humans and agents. Google Drive, Gmail, Calendar, and every Workspace API. 40+ agent skills included.

Media 1
❀️14,232
likes
πŸ”1,490
retweets
πŸ–ΌοΈ Media
N
nanbeige
@nanbeige
πŸ“…
Mar 05, 2026
159d ago
πŸ†”30220863

In both LeetCode's Weekly Contests (Weekly Contests 489–491) and the HMMT February 2026 (Harvard-MIT Mathematics Tournament), Nanbeige4.1-3B's performance not only significantly outperformed that of Qwen3.5-4B but also surpassed Qwen3.5-9B. https://t.co/2guwzB3yNa

Media 1
πŸ–ΌοΈ Media
N
nanbeige
@nanbeige
πŸ“…
Mar 05, 2026
159d ago
πŸ†”30220863

In both LeetCode's Weekly Contests (Weekly Contests 489–491) and the HMMT February 2026 (Harvard-MIT Mathematics Tournament), Nanbeige4.1-3B's performance not only significantly outperformed that of Qwen3.5-4B but also surpassed Qwen3.5-9B. https://t.co/2guwzB3yNa

Media 1
πŸ–ΌοΈ Media
A
AcerFur
@AcerFur
πŸ“…
Mar 05, 2026
159d ago
πŸ†”13955357

Also, come on OpenAI. If you want an automated AI researcher, this needs to start going up, not down. https://t.co/0ZQ4UhdNyu

Media 1
πŸ–ΌοΈ Media
πŸ”jeremyphoward retweeted
A
Acer
@AcerFur
πŸ“…
Mar 05, 2026
159d ago
πŸ†”13955357

Also, come on OpenAI. If you want an automated AI researcher, this needs to start going up, not down. https://t.co/0ZQ4UhdNyu

Media 1
❀️522
likes
πŸ”27
retweets
πŸ–ΌοΈ Media
πŸ”jeremyphoward retweeted
A
Acer
@AcerFur
πŸ“…
Mar 05, 2026
159d ago
πŸ†”13955357

Also, come on OpenAI. If you want an automated AI researcher, this needs to start going up, not down. https://t.co/0ZQ4UhdNyu

Media 1
❀️522
likes
πŸ”27
retweets
πŸ–ΌοΈ Media
← PreviousPage 694 of 1101Next β†’