Your curated collection of saved posts and media
Imagine every pixel on your screen, streamed live directly from a model. No HTML, no layout engine, no code. Just exactly what you want to see. @eddiejiao_obj, @drewocarr and I built a prototype to see how this could actually work, and set out to make it real. We're calling it Flipbook. (1/5)
New work with @AlecRad and @DavidDuvenaud: Have you ever dreamed of talking to someone from the past? Introducing talkie, a 13B model trained only on pre-1931 text. Vintage models should help us to understand how LMs generalize (e.g., can we teach talkie to code?). Thread: https://t.co/usVP9KErdv
@VBarsoum βLove > Logicβ :D https://t.co/5HyIjHxrMz
Here is a 2nd batch of April architecture drops. What a month! - Ant Ling 2.6 1T - Minimax M2.7 - Xiaomi MiMo V2.5 - Poolside Laguna XS.2 - Tencent Hy3-preview - IBM Granite 4.1 https://t.co/ILjjyZYeC5
As always, more details and higher-res versions at https://t.co/NO7z6XSRHS
@giffmana I used to get one every year for backups https://t.co/N2PxYG93Jv
Back from a little family break! Lots has happened, and Iβm planning to do a deeper dive into the most interesting architectural components (soon). Btw, are there any major architectures I missed below? https://t.co/HEhXUjxaDY
Come celebrate after Google IO with the DeepMind team in SF on May 21st :) Lots to look forward to!! Link in π§΅ https://t.co/FP6y97CJmc
Have open source models closed the gap with proprietary ones? We've tracked three years of Arena data across three arenas. The short answer: mostly yes. In Text Arena, the proprietary winner had a +250 Arena lead. By early 2025, it had fallen to low double digits, and at its narrowest was almost closed entirely. Today, the proprietary lead is about +30 points. It separates #1 from roughly #18 on the current leaderboard. - Open source has quickly closed most of the gap - The biggest gains happened before 2025 - The remaining gap is small in points, but still large in rank Get a deeper look into the race for Code Arena: Frontend and Expert prompts in the thread π§΅
Feeding the hungry hungry hackers at the @aiDotEngineer SG hackathon with @unprofeshme π₯π₯π₯ https://t.co/J51RisTCjG
Feeding the hungry hungry hackers at the @aiDotEngineer SG hackathon with @unprofeshme π₯π₯π₯ https://t.co/J51RisTCjG
We are so incredibly honoured to welcome Singaporeβs Minister for Foreign Affairs, Dr Vivian Balakrishnan, to the lineup of speakers for @aiDotEngineer Singapore. A few weeks ago, Minister Balakrishnan casually dropped a technical writeup of his personal AI system online. Raspberry Pi. Claude. Local embeddings. Knowledge graphs. A full architecture breakdown. And the global AI community noticed because it reflected something bigger: a willingness to engage with these systems directly, publicly, and practically. To kick off AIE Singapore, Minister Balakrishnan will share his experience experimenting with open-source AI tools and building a βsecond brainβ workflow, alongside broader reflections on how AI may reshape global dynamics, and the way people work, think, and manage information. In a role that demands navigating enormous volumes of information and constant context-switching, his reflections will set the tone for the conference in exactly the way we hoped: That meaningful conversations about AI should not stay abstract. They should involve understanding its parameters through practical engagement with the technology. And that Singapore has become the place where that kind of engagement happens seriously.
Og breakfast is back Get an egg, dip that bread in it and enjoy a great meal. Also total cost was 6 USD hahahahaha https://t.co/hBHMM7U3Lo
When u forget to bring the right shoes to Singapore π https://t.co/hA25NmBWPZ
Yesterday Fitbit Air launched, but did you know it comes with a new @googlehealth API? You can build AI agents, MCP servers, or CLIs on top of your sleep and heat data. - 31 different data points from exercise to sleep, heart rate, or SpO2. - Webhooks push real-time notifications when health data changes. - Support Read or write data, request only the permissions you need. - Query by time range, roll up daily summaries, or paginate results. I am whoop guy, but might be good reason to explore this. Full getting started codelab below.ππ»
in sg for aie, building agents/robots? let's meet https://t.co/91y1hp9ZyG
in sg for aie, building agents/robots? let's meet https://t.co/91y1hp9ZyG
we're so excited that @Zai_org is joining ai engineer singapore as a diamond sponsor π tsinghua university spin-off, first major chinese llm company to ipo ($31B valuation), and glm-5.1 just dropped as open-source under MIT license. we're thrilled they are bringing three sessions to ai engineer singapore: π€ main stage - @ZixuanLi_ (head of https://t.co/CxhkRbzAq2) on "glm-5.1: towards long-horizon tasks" π§ leadership track - "https://t.co/CxhkRbzAq2: more than the creator of glm" π οΈ workshop - "building in the open with https://t.co/CxhkRbzAq2" with wenyu li & xiaopu peng. your data, your stack, your call. see what 8-hour autonomous building looks like in practice, then get hands-on and build something yourself. may 15β17 Β· capitol kempinski, singapore
8 days until Google I/O β³ Calm on the surface...mostly π΅π΄π‘π’ https://t.co/ytIpfTraXd
8 days until Google I/O β³ Calm on the surface...mostly π΅π΄π‘π’ https://t.co/ytIpfTraXd
QoL: Every Gemini API documentation page now has a interactions API version, readable for humans or agents via markdown https://t.co/Sn03tQBlsm
Gemma 4 is a highly-capable model in Japanese! Amazing to see great results for it on the latest Swallow Leaderboard v2 results. β¨ Excited to see whatβs next in the open model space! https://t.co/aK4raCKIeh
GPT-5.5 takes over WolfBench! Itβs now the #1 model, ahead of Claude Opus 4.7 and 4.6, GPT-5.4, Sonnet 4.6, Kimi K2.6, Gemini 3.1 Pro, and more. Notable findings after 30 runs (40h runtime, >1.7B tokens, ~$3K cost): - @OpenAI's GPT-5.5 is the best model we ever tested. - @cursor_ai's Agent CLI (CA) is the best agent we ever tested. - @NousResearch's Hermes Agent (HA) outperformed OpenClaw (OC). - With Hermes, going from medium to xhigh reasoning only improved consistency, not capability. Note: This is WolfBench, where we look at more than just the average score, because one metric is not enough. The golden β score is the actual 5-run average, which most other benchmarks report as their only score. β shows the ceiling (what percentage of the full benchmark this model+agent combination solved at least once across all runs). β shows the solid base (what percentage of the full benchmark it solved consistently in every run).
NEW: @IBM Granite 4.1 8B is live on W&B Inference! $0.05 / $0.10 per 1M tokens. 131k context. Apache 2.0. Build production agents with native tool calling, trace every call with @weave_wb, and run it all on @CoreWeave's platinum-grade AI cloud. π§΅ https://t.co/bbFxtUiO1H
W&B Inference provides frontier and open-weights models, observability included, on @SemiAnalysis_ platinum-grade infrastructure. Day-0 launches, every time. Try it below! https://t.co/TfJWyVu1KU
Stop hunting through config for the same 3 links every run. Pin them as References at the top of the W&B Overview tab. Renders as markdown, sits above Notes and Tags, always one click away. https://t.co/WUGTsIBvZR
Can confirm, @cursor_ai is the best harness we've tested on @WolfBenchAI so far! @WolframRvnwlf tests Harness x Model, and Cursor (before the SDK) is the best one we've ever tested! https://t.co/YVorHrBZfD
lol Cursor is a better harness for both GPT 5.5 in Codex AND Opus 4.7 in Claude Code how is that possible?! https://t.co/oNALWpjTtS
Production VLA pipelines scale three workloads across three hardware profiles, plus thousands of stochastic sim rollouts. Join @anyscalecompute + @wandb May 12 to orchestrate it end-to-end with Ray. Register: https://t.co/6vr1VsTD9d https://t.co/uEwZmwcK8L

Not too long ago, this would've been a weeks-long team project. Now? One developer, one notebook: an app that lets you upload a photo of your room and swap any lamp in it with one from a store catalog. The prototype-to-prod gap is closing fast. π https://t.co/cMJdZyd3IN
π¨ @GeminiApp for image generation π @weave_wb for observability + evals across models (accuracy, latency, cost) π @marimo_io notebook for the prototype, then shipped as a web app Blog π https://t.co/7aBN5d0vXb
AI is not going to replace human beings,β said @mcannonbrookes speaking to @l2k on @wandbβs podcast. And it definitely stayed with me. AI & agents are only going to enable us to do more, build more, and help humanity realize its full potential. #TLVCPartner https://t.co/dDt51kZgqB
Kimi K2.6 from @Kimi_Moonshot is purpose-built for coding agents. As of today, CoreWeave ranked highest in @ArtificialAnlysβs inference benchmark on Speed vs. Price for K2.6. Speed, scale, and economics. All three at production grade. https://t.co/LyAe4ScapY