Your curated collection of saved posts and media
LTX Ripple is a new efficient IC LoRA approach for video editing โ๏ธ๐๏ธ edit just the first frame and have the edits ripple into the rest of the video. lands perfect edits, incredibly fast! LTX 2.5 based โถ๏ธ https://t.co/ryTutBeiOc https://t.co/RJSsqwH7mH
Louen Pottierๆฐ้็บใฎ็ฉ็AIใLaGSplatใ ๐นในใใๅ็ป1ๆฌใใ็ฉ็ๆณๅ๏ผใจใใซใฎใผๆฃ้ธใชใฉ๏ผใ่ชๅๅญฆ็ฟ ๐นไบๅใซๅใ่จๆธฌใใชใใฆใใๅ็ปๅ ใฎ็ฉไฝใใๆผใใๆไฝใๅฏ่ฝ ๐น็พๅฎใฎๆ ๅใใ็ดๆฅ็ฉ็ใทใใฅใฌใผใฟใไฝใใญใใใ้็บใฎๅฟ็จใธใฎๆๅพ Webใใฉใฆใถใงไฝ้จๅฏ่ฝใๆ ๅ ฑๅ ใฏใชใๆฌใ https://t.co/KpgNNIGN8m
Regarding Sentence Transformers v6.0 update: A dense model compresses a whole text into one vector, then compares two vectors. A multi-vector model keeps one vector per token, scores every query token against every document token, takes the best match for each, and sums those. https://t.co/ksnigtsLpm
Reminds me of novel locomotion policies discovered in MuJoCo environments, except that they work in the real world!
Slow-motion look at the 400m championโs running form at the World Humanoid Robot Games. https://t.co/Yt2Vp8DLcx
Single video in, 4D human out. 4DAnyone turns a casual monocular video into a 4DGS model. No camera rig, no calibration, no tripod. Project: https://t.co/hIArgF0YmQ https://t.co/JB3S9bqggd
one real issue I have with various agents now is "where is this code going to run" - does it need access to my local filesystem - can I fire it off from my phone and if so does it run on the cloud or needs my Mac mini fired up / synced? The current UX metaphors around this seem to need work.
๐ Best Resource Paper Award at #ACL2026 @aclmeeting ๐Thrilled to share that our deep research benchmark paper "HSCodeComp: A Realistic and Expert-level Agent Benchmark for Hierarchical Rule Application" reveived a Best Resource Paper Award #ACL2026! ๐ฏ Hoping this sparks more focus on reliable AI Agents in the real world. https://t.co/ToFSPPaYb3

How to approach an unknown language in the ocean? Hereโs one of the first cases of AI interpretability leading to a scientific discovery -- in whales. We built an artificial baby model that learns language directly from raw sound and trained it to imitate whale speech. Then we looked inside. Our interpretability method recovered the properties biologists already thought were meaningful and pointed out those that had not been considered before. This was the initial clue that eventually led to the discovery of vowels in sperm whales. Understanding AI and reframing language as informative imagitation can help us step outside our human biases and discover new realities about the natural world. Published in Royal Society Open Science.
FastAPI Conf speaker news! ๐ข @jxnlco, creator of Instructor, the person that convinced AI Labs to provide structured output with Pydantic (automatic compatibility with FastAPI). He's bringing us some insights directly from the Codex team at @OpenAI ๐ค https://t.co/ioc3vjSHvX https://t.co/C1zlrff3aX
This is what faster-than-real-time video generation looks like.
A line in AI video was crossed, in my experiments with just the web interface, H3 Max can now create reasonably high quality AI video in less time than it takes you to watch it. This is realtime from the moment I pushed the "generate" button (and also includes prompt enhancement)
This paper does a good job of showing the promise and gaps of autonomous AI scientists. Big question is how much more advanced models close those gaps.
Today, we're excited to share early progress on using Gemini to accelerate scientific discovery in the real-world. We present an extension of Co-Scientist which we use to collaborate with scientists across materials science, biology, and computer science. https://t.co/jwDF5aSK28
๐ค Debug Your Scripts in VS Code with GitHub Copilot Tired of remembering all the syntax for a launch.json file? In this quick tip, learn how to use GitHub Copilot to analyze your package.json, generate the right debug configurations, and launch your scripts directly with the built-in VS Code debugger.
Video world models shouldn't just render plausible pixels, they should understand "how the world evolves". ๐ค Introducing Latent Dynamics Reasoning (LDR). Instead of predicting future frames directly, LDR maps past frames into structured latent states, then rolls those states forward through kinematic integration. It learns only the higher-order motion residuals that drive the rollout. ๐ To our knowledge, LDR is the first video world model to extrapolate learned dynamics beyond its training distribution. ๐ช Paper / code / models / data are now public, check them out! ๐ฅณ - Paper: https://t.co/Mnh1sD9jLf - Code: https://t.co/uJhu28darp - Model: https://t.co/1fJLmwFwTL - Data: https://t.co/Pxl0eirIgo

Introducing Workflow -- make your AI pipeline the interface. A drag-and-drop canvas where every node runs on its own, every intermediate output stays visible, deploys to Spaces in one command, and the whole graph is also a REST API. Guide: https://t.co/6MM0wJZxMX
One of the trickiest problems you run into with voice AI in the real world is that people often call in messy environments. There might be other voices in the background, or the person might be talking to other people at the same time. Super cool work from Decagon Labs on how we've tackled this problem!
Conversational voice agents need to know when a different person is speaking, but not every background voice should count. We combined speaker embeddings with a post-trained audio-language model to determine when a speaker change matters. https://t.co/Cg5HKtyaUf
Breeze TTS 2 by @BreezeBlueX dropped on Hugging Face as the #1 open weights TTS model on @ArtificialAnlys It does ๐จ Voice Design (prompt a voice) ๐๏ธ Voice Reference (upload a voice) ๐๏ธ Voice Direction (upload & prompt a voice) vibe it on Spaces โถ๏ธ https://t.co/MCKL1l4y6R https://t.co/QTgLO50JtD
Breeze TTS 2 is now the leading Open Weights TTS model in the Artificial Analysis Provider Voices Speech Arena, surpassing Fish Audio S2 Pro by 90 Elo points Breeze TTS 2 is the latest TTS model from @BreezeBlueX, supporting 50 languages, voice generation from text prompts, and
The Perplexity team and @NaderLikeLadder from NVIDIA demoed Computer for me during my lunch break at Hot Chips. This makes DGX Spark a genuinely easy-to-use, useful agentic workstation. Excited to plug it into my own workflows and see what it can do. (Excuse the etched hat it was sunny ๐คฃ) Speed of light $NVDA
sat down and actually read @rosmine's research paper on this. https://t.co/vxqTirXxrS first of all, i was guilty of having a "take" based on this post, pointing out how the copy in the screenshot is not a great example of what "good writing" is, despite scoring 100 on "human written" in @pangram. I recommend actually reading the work before commenting on it - because sloppy takes are as lazy as sloppy texts. (shame on me for joining the band wagon) but sleeping on it, i am grateful that @rosmine took the time to dive into this stuff. AI-slop fatigue is real, and we should support efforts to make agents produce communication that is clear and lucid, and doesn't feel overly synthetic but there are some assumptions in this, however, that is interesting to question when it comes to "what makes for good writing" as far as i understand, Deft is trained to have more diversity/variation in tokens, so less repetition of what we are recognizing as "AI-tells" (certain word and stylistic choices). And i totally agree with @rosmine that: "LLMs are not the cause of slop. Lack of effort/care is. If you spend days researching and planning a blog post, and put all the information into a detailed, well-structured outline, and ask ChatGPT to generate the post based on the outline, then the output will be interesting to read, even if the text has a lot of em-dashes." I argued the same in our eng blog announcement post yesterday: https://t.co/sV3cEP4S9G BUT! I still feel that this report (at least somewhat), but especially the various takes on it, conflates something sounding "human" with it being "good." Tricking @pangram doesn't make a text well written. Some reflections: - Making writing sound more "human" by means of adding more variation in word/style choices, doesn't make it better - What makes for a "good" text is highly contextual. If you are writing a recipe or instructions, repetition and stylistic stringency is highly desirable - AI-tells cuts deeper than just word choices, observant readers will start to be sensitive to the lack of certain devices, ways of arguing, structure, etc. - Interesting writing often comes from doing synthesis of unexpected things in a way that bring clarity. I have still to see that from LLMs as we usually interact with them (but it's probably possible to get them to do this). with that being said - i'm grateful that this research was shared with the wider community, and will applaud any effort to make the user experience of interacting with agents, and the stuff we make with agents, better. ๐ซก
Announcing Deft, a new AI lab for better writing, cofounded with @jmrphy See the picture for launch announcement the Deft model wrote for itself Currently, 86% of user queries are fully human according to pangram. This is still a small beta model and it might make mistakes. We

The Perplexity team and @NaderLikeLadder from NVIDIA demoed Computer for me during my lunch break at Hot Chips. This makes DGX Spark a genuinely easy-to-use, useful agentic workstation. Excited to plug it into my own workflows and see what it can do. (Excuse the etched hat it was sunny ๐คฃ) Speed of light $NVDA
2 months ago, we crashed Jensen's board meeting to show him Perplexity running locally on DGX Spark ๐คฃ (lol im holding architecture diagrams like he's gonna look at em during a board meeting) Local AI used to be for enthusiasts. Folks were running tiny quantized models on underp
Modular's price-performance on @Zai_org's GLM-5.2 (Non-reasoning) lands right on the Pareto frontier in @ArtificialAnlys' latest benchmark: near-top speed without the near-top price tag. We're just getting started, and we're ready for GLM-5.3. Expect to see a lot more incredible results. ๐
โThe fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.โ Hereโs my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them to audit the evals I built for my creator skills live. They then demoed a free skill that you can use in Claude Code or Codex to build reusable evals from your feedback. Some quotes from Shreya and Hamel: โBottom-up evals come from looking at lots of sample outputs and turning that into eval criteria. AI is very bad at coming up with them. Thatโs all you.โ โThe agentโs job is not to invent new feedback. But it can help you group and distill the feedback into actionable rubric criteria.โ โAll your competitors can point Claude at their product and say, โFind all the errors.โ What matters is how much taste you can infuse beyond that.โ ๐ Watch now: https://t.co/BuTklgfbr4 Thanks to our sponsors: @WisprFlow: 4x faster than typing with your voice https://t.co/oqHJ8bN3ll @linear: The AI agent platform for modern teams https://t.co/tgWf9oL4bs
AI2 models ๐ค EleutherAI interp tooling Open science
AI2 models ๐ค EleutherAI interp tooling Open science
What kinds of training data shape different model capabilities? A @GeorgiaTech team used our fully open model flow to trace Olmoโs performance on social/general reasoning and social-science/STEM knowledge tests back to the types of text it trained on. ๐งต https://t.co/qrPoTrdj2B
Introducing https://t.co/SLQRCKIF1x Deploy your agent as a function. Get a computer for it. - a real Linux machine per session - ffmpeg, chromium, git, install anything - durable: hibernate + resume - your keys never enter the runtime The cloud of the last decade was built for apps. The next one is for trillions of always on agents. @opencomputerhq is their home.
The Perplexity Agent API now gives developers access to 41 frontier models across 9 providers in one endpoint. Build multi-model agent workflows with built-in tools like web search, finance search, fetch, and sandboxed code execution. https://t.co/MZfzN0EcET
Granite 4.2 8B is live on CoreWeave Serverless Inference on day 0. Reasons natively, calls tools on its own and you do not have to worry about infra to manage. @IBMResearch built it for enterprise agents. We keep it served. https://t.co/RHIkXibUGw
Meet Granite 4.2, IBMโs latest family of open models purpose-built for enterprise agentic AI. With new native reasoning capabilities, Granite 4.2 can plan, reason, self-correct and reliably use tools to help automate complex enterprise workflows โฌ๏ธ
I still can't wrap my brain around the scale of global inference. I guestimate we're at around 10 QUADRILLION tokens per month now? I wonder what fraction ever get read.
Codex has done 40 trillion tokens since launch (3 weeks ish). I'm trying to down-convert that into some kind of unit I can understand. I think I write <1M tokens/year. So a million lifetimes worth? For 50M devs, 800k tokens per developer on earth? ~20M tokens per second?
The tasks are taken from real-world engineering workflows and verifiers are authored by human experts It's the beginning of a broader effort at Seldon to create evals and RL environments advancing computer-use agents https://t.co/qqsf4y0srO
Take your pick of interactive experiences like: ๐ค Orchestrating specialized agents with GitHub Copilot CLI โ๏ธ Building agentic workflows that automate repo tasks โ๏ธ Going from idea to merged PR with the GitHub Copilot app ๐ก๏ธ Automating quality signals and catching issues earlier
๐จ Claude Fable 5.1 is MONSTROUS. Anthropic just dropped a new frontier model that is crushing Fable 5 across coding, reasoning, computer use and agentic work. It scores: 55.8% Terminal-Bench 73.4% CursorBench 60.9% Humanityโs Last Exam 77.9% OSWorld 31.4% AutomationBench 1853 GDPval-AA v2 And somehow, itโs cheaper. $10/M input $50/M output $0.25/M cache reads 75% cheaper. It can also run long coding and research workflows unattended for hours, while keeping track of what itโs doing and verifying its own work.
๐ Looking for something from earlier in a conversation? Prompts that changed files also show the number of lines added and removed, with the option to review those changes directly. https://t.co/dZyEgkKVg1
Most extraction tools treat spreadsheets like PDFs. They flatten the file into text or markdown, then ask a model to infer the original structure. But spreadsheets depend on structure. Headers, formulas, merged cells, and hidden rows give every value its context. Strip that away, and you map the right number to the wrong metric or period. That's why we built native spreadsheet extraction into the LlamaParse platform. Instead of flattening your workbook to text, it reads the raw cells directly and maps the data to your schema. Available today in beta on the agentic_plus tier. Give it a spin on your messiest .xlsx, .xls, or .csv files. Docs: https://t.co/z8l3MqRpDj