Your curated collection of saved posts and media
@KenWattana Yeah, agree that it's a hard problem. It might be the EQ version of uncanny valley. https://t.co/7zImchwKMo
I donβt know exactly whatβs going on here, but it does feel AI-related. Unlike PM and eng, which started growing in 2024 (two years post-ChatGPT), design didnβt. If I had to venture a theory, Iβd say that because AI is allowing engineers to move so quickly, thereβs less opportunityβand less desireβto involve the traditional design process. That said, youβd think design would become a differentiator as more products compete for attention. Something to think about for your company! Weβll keep watching this trend and AIβs impact on org design more generally. One interesting observation we made when we went a level deeper: the ratio of demand for PMs vs. designers has flipped. In mid-2023, we went from more open designer roles to more open PM roles. And ever since, PM demand has been pulling away (currently 1.27x). This will be another trend to monitor, in terms of how AI is reshaping org design.
turns out Apple is ~distilling~ Google's Gemini model to produce other AI models for the Siri/consumer features it wants to launch https://t.co/cXjFxzfk7d
The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature Nature: https://t.co/nNfpSV5e5I Blog: https://t.co/i6h8LVQOdl When we first introduced The AI Scientist, we shared an ambitious vision of an agent powered by foundation models capable of executing the entire machine learning research lifecycle. From inventing ideas and writing code to executing experiments and drafting the manuscript, the system demonstrated that end-to-end automation of the scientific process is possible. Soon after, we shared a historic update: the improved AI Scientist-v2 produced the first fully AI-generated paper to pass a rigorous human peer-review process. Today, we are happy to announce that βThe AI Scientist: Towards Fully Automated AI Research,β our paper describing all of this work, along with fresh new insights, has been published in @Nature! This Nature publication consolidates these milestones and details the underlying foundation model orchestration. It also introduces our Automated Reviewer, which matches human review judgments and actually exceeds standard inter-human agreement. Crucially, by using this reviewer to grade papers generated by different foundation models, we discovered a clear scaling law of science. As the underlying foundation models improve, the quality of the generated scientific papers increases correspondingly. This implies that as compute costs decrease and model capabilities continue to exponentially increase, future versions of The AI Scientist will be substantially more capable. Building upon our previous open-source releases (https://t.co/H1tBT14Yx8), this open-access Nature publication comprehensively details our system's architecture, outlines several new scaling results, and discusses the promise and challenges of AI-generated science. This substantial milestone is the result of a close and fruitful collaboration between researchers at Sakana AI, the University of British Columbia (UBC) and the Vector Institute, and the University of Oxford. Congrats to the team! @_chris_lu_ @cong_ml @RobertTLange @_yutaroyamada @shengranhu @j_foerst @hardmaru @jeffclune
The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature Nature: https://t.co/nNfpSV5e5I Blog: https://t.co/i6h8LVQOdl When we first introduced The AI Scientist, we shared an ambitious vision of an agent powered by foundation models capable of executing the entire machine learning research lifecycle. From inventing ideas and writing code to executing experiments and drafting the manuscript, the system demonstrated that end-to-end automation of the scientific process is possible. Soon after, we shared a historic update: the improved AI Scientist-v2 produced the first fully AI-generated paper to pass a rigorous human peer-review process. Today, we are happy to announce that βThe AI Scientist: Towards Fully Automated AI Research,β our paper describing all of this work, along with fresh new insights, has been published in @Nature! This Nature publication consolidates these milestones and details the underlying foundation model orchestration. It also introduces our Automated Reviewer, which matches human review judgments and actually exceeds standard inter-human agreement. Crucially, by using this reviewer to grade papers generated by different foundation models, we discovered a clear scaling law of science. As the underlying foundation models improve, the quality of the generated scientific papers increases correspondingly. This implies that as compute costs decrease and model capabilities continue to exponentially increase, future versions of The AI Scientist will be substantially more capable. Building upon our previous open-source releases (https://t.co/H1tBT14Yx8), this open-access Nature publication comprehensively details our system's architecture, outlines several new scaling results, and discusses the promise and challenges of AI-generated science. This substantial milestone is the result of a close and fruitful collaboration between researchers at Sakana AI, the University of British Columbia (UBC) and the Vector Institute, and the University of Oxford. Congrats to the team! @_chris_lu_ @cong_ml @RobertTLange @_yutaroyamada @shengranhu @j_foerst @hardmaru @jeffclune

This week on Clear+Vivid @AlanAlda speaks with Gary Marcus. He has been a relentless critic of the extravagant claims made for the current generation of AI based on Large Language Models, or LLMs... π https://t.co/U9UkMc2CuN https://t.co/NJUlH5deKr
Better AI may create a new kind of problem. As systems become more reliable, people pay less attention, making oversight weaker rather than stronger. When things mostly work, vigilance tends to drop. The risk is subtle. The more we trust AI, the easier it becomes to stop checking it. https://t.co/mV06aB7emN @whartonknows
So proud to see F.03 make history as the first humanoid robot in the White House π€ πΊπΈ https://t.co/tXsxpEErsi
So proud to see F.03 make history as the first humanoid robot in the White House π€ πΊπΈ https://t.co/tXsxpEErsi
Minicor (@minicor_) builds self-healing desktop automations for AI companies whose customers run on legacy desktop software with no APIs. Congrats on the launch, @faizchishtie and @sahee_d! https://t.co/mAYfH8KMoS https://t.co/wUZTU9ehni
MinerU-Diffusion Rethinking Document OCR as Inverse Rendering via Diffusion Decoding paper: https://t.co/nLWiJNSdsQ https://t.co/PoVYI63Yt7

WildWorld A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG paper: https://t.co/u0btopls0r https://t.co/qKyFgk3dwT
SpecEyes Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning paper: https://t.co/LkFvuVgs1v https://t.co/aKwrEYnPNH

the most influential piece of writing wisdom for me is ursula k. le guin, who wrote a dozen books about guys sailing, saying: βYeah idk, i've never been in a boat. I was just sort of guessingβ https://t.co/HzDW57Debm
this is how fantasy should operate. the physical/temporal feel unreal but are grounded by the humanity. I think too often fantasy authors opt for the reverse
Melania Trump enters the room with a robot https://t.co/l9ljIYby0E
clicking refresh everyday, m2.7 on HF soon: https://t.co/wPHHa7hIsw
Thanks @TheAhmadOsman, Iβve started preparing the model card of the HF β M2.7 coming soon!
Word docs are one of the most common file formats people process in LlamaParse, and they've always been surprisingly frustrating to parse well. Here's the counterintuitive part: .docx actually has better structural information than most document formats. We just haven't been able to fully use it. Until now. A .docx file is a ZIP archive of XML files. That XML knows everything: cell boundaries, merged cells, column and row spans, nested tables, formatting tags, hyperlinks. A PDF of the same table has none of that. It's just text positioned at coordinates and line intersections that a parser has to reverse-engineer into structure. The hard part with Word XML isn't extracting the table content. It's knowing which page it's on. Word is a flow format β there are no page boundaries in the XML. Pagination depends on the renderer, fonts, margins, line-height. The same .docx renders differently in Word, LibreOffice, and Google Docs. We built a technique to resolve this, mapping Word XML table elements to their correct page positions in the rendered output. We now get the original document structure AND know exactly where each table appears. The quality improvement is most significant for: Β· Tables with rich cell formatting (bold, italic, strikethrough, superscript, lists inside cells) Β· Merged cells and column/row spans Β· Nested tables (tables inside table cells) If you're processing Word docs with table-heavy content, try it out. π Full writeup: https://t.co/aAEFkvfycG

π Signup to LlamaParse to try it out: https://t.co/QQzVCOwiVl
I had access to the new Google Lyria 3 Pro music AI. Its quite good. I've been ruining(?) Rilke by giving the AI the First Elegy & asking it to make it "more 1990s boy band" ("oooo the beginning of terror, girl") Catchy! It is also nuts that you can ask an AI to do this & it can https://t.co/xQIbO6XOUL
Will Google solve disease in 10 years? Not without data. We can commoditize drug design, and STILL not solve most diseases. You need @precigenetic. https://t.co/TAckLIU7rg
β‘οΈ New on Lovart: Move Object β Select any object with rectangular or lasso tool β Move it wherever you want β Prompt optional modifications β One clean, consistent image No masks. No layers. No re-roll.
We're launching RunClaw to kill OpenClaw. OpenClaw costs $700 to set up. RunClaw costs $1 no setup. > OpenClaw canβt build you a website > Canβt generate a video > Canβt make a slide deck > Has 9 security CVEs RunClaw does all of it. Better agents. More secure. Always in your DMs. Try it now for $1.
Last month, we released Lyria 3, enabling you to create tracks with lyrics from text, image, or video prompts. Now, weβre introducing Lyria 3 Pro, which expands upon our music generation model to offer additional advanced capabilities. Whatβs really special about this upgrade is that the model now understands the architecture of music. This makes it possible to prompt for intros, verses, choruses and bridges + generate songs with more complex transitions. You can also create tracks up to 3 minutes long, a big change from previous models that were limited to 30 second tracks. Use Lyria 3 Pro to build upon your existing creativity. Weβre excited for your beats to drop πΆ
Bro, that shit you guys are hyping dropped in April last year. Why are you acting like itβs new now? https://t.co/7vdZ34UVmL
This is one of the most interesting papers on self-improving agents for this year. (bookmark this one) Most self-improving AI systems hit the same wall: the mechanism that generates improvements is fixed and can't improve itself. This new work from Meta and collaborators breaks through this limitation. They introduce Hyperagents, self-referential agents where the self-improvement process itself is editable. The DGM-Hyperagent combines a task agent and a meta agent into a single modifiable program, enabling metacognitive self-modification. It autonomously discovers innovations like persistent memory and performance tracking, and these meta-improvements transfer across domains and compound across runs. Why does it matter? - On paper review, DGM-H improved from 0.0 to 0.710 test accuracy. - On robotics reward design, it went from 0.060 to 0.372. - Transfer hyperagents achieved 0.630 on Olympiad-level math grading in a domain they were never trained on. This is a step toward AI systems that don't just find better solutions but continuously improve how they search for improvements. Paper: https://t.co/Q0f7zWhNMD Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX

Building robust phone agents used to require deep expertise in the nuances of Voice AI. Not anymore. Introducing Norm. Just tell Norm what kind of phone agent you want to build and it does it all for you. Sign up and try it β https://t.co/aeM7IFgSzk https://t.co/QsbdsYNq6H
Nice cheat sheet for Claude Code. https://t.co/ikGzbSqjRK
Nice cheat sheet for Claude Code. https://t.co/ikGzbSqjRK
We got our hands on Reachy from @pollenrobotics and @huggingface and we're already deep into experimenting with it. So good to see more open-source players making robotics affordable and accessible. #robotics #opensource #huggingface #pollenrobotics https://t.co/Bw3Q0SpqDH
Introducing the first discrete diffusion pipeline for text in Diffusers -- LLaDA2 by @TheInclusionAI π₯ It follows an MoE architecture w/ 16B total params. It is definitely not SOTA across the board, but it hopefully flips that soon. Check out the links below to know more β¬οΈ https://t.co/dWN9R1O0RN
Little quality of life improvement in the Codex app. You can now search your threads for faster navigation. And if you don't want to take your hands off the keyboard you can open it directly using Cmd+K https://t.co/TPbaNccUVn
π₯³Excited to see that π§¬Xperience-10M𧬠has been listed in @huggingface "most downloads datasets", hitting *1M downloads within 1 week*! - Dataset link @HuggingModels: https://t.co/MiKukkLLOv https://t.co/9hEVInwXuH
Today Ropedia releases Xperience-10M at #GTC day 1 β World largest real human 4D interaction dataset at 10M scale. Each trajectory aligns: β’ visual observations β’ spatial structure β’ human motion β’ interaction dynamics β’ task semantics A new foundation for physical and
