Your curated collection of saved posts and media
90% reduction in response latency. That's what enterprise AI should look like. @AdaptiveML deployed a Llama-powered AI agent on AWS to transform patient service operations for CCS, making chronic care support faster and more reliable than ever. https://t.co/cHXrHAA4AN
We are honored to announce the Test of Time awards for #ICLR2026 π This award recognizes papers published 10 years ago at ICLR 2016 that have had a lasting impact on the field: https://t.co/JqYiqrAvgz https://t.co/xIZGtUPAB4
Who decides what counts as progress in AI, Ethics, and Society @AIESConf? Nominations are open for (Senior) Program Committee members for the 9th AAAI/ACM Conference on AI, Ethics, and Society, to be in MalmΓΆ, Sweden, October 12β14, 2026. Shape the field! https://t.co/eOaekzEIhe https://t.co/EQobqycq02
Computer use with any model Hermes Agent Γ @trycua https://t.co/z2iQHmj1Q8
April's community meeting is happening tomorrow, and we're spotlighting three impressive projects: - Marrow, an Apache Arrow implementation in Mojo by @kszucs_ - Mojo support on @tensarahq - MAV ffmpeg Mojo bindings Join us via Zoom: https://t.co/tUlQ7LmnCG
Two days left until @AMD's AI DevDay! Don't miss @clattner_llvm's luminary talk covering how we fused FLUX.2 into a single execution graph, with a 3.8x speedup torch.compile on MI355X, under 3.5s per image, sub-700MB container. Plus, 5.5x lower cost than running on competing hardware. Stop by the Modular booth to connect with our team, learn about open roles, and score swag. Grab your spot: https://t.co/Pa1e36BTZn
Ling-2.6-flash is now officially open-sourced! A fast, token-efficient Instruct model built for real-world agent workflows. 104B total parameters Β· 7.4B active parameters Available in BF16, FP8, and INT4 variants for different deployment needs. Key strengths: - Fast generation: 215 tokens/s on Artificial Analysis Output Speed - High token efficiency: only 15M tokens on the full AA Intelligence Index evaluation - Real task execution: strong performance across coding, document processing, and lightweight agent workflows - Improved experience: better Chinese-English switching and smoother compatibility with mainstream coding frameworks
Tomorrow, @clattner_llvm takes the stage at @AMD's AI DevDay. He's covering FLUX.2 on the MI355X with Modular: 3.8x faster than torch.compile, 1024x1024 images in under 3.5 seconds, deployment container under 700MB. Register for DevDay: https://t.co/pEiBqEzaXR
We're on the floor at @AMD AI DevDay! Stop by our booth to talk high-performance inference with AMD + Modular, and don't miss @clattner_llvm's luminary talk at 3:10 PM. https://t.co/exQ5GULe1D
Highlights from @AMD AI DevDay π· Great to see so many developers at the booth and our reception the night before, and even better to watch their reactions when they saw MAX and Mojo π₯ in action. https://t.co/UvoRgxJNJQ

The changelogs are the best place to explore all of 26.3's improvements. π MAX: https://t.co/8E0d3gDgjE π Mojo: https://t.co/bjtOXpq2p3 Install or upgrade by running `uv pip install --upgrade modular`. Tell us what you're building with 26.3: https://t.co/QQIVUie2kQ
Mojo π₯ 1.0 beta is out! Now we want to hear from you. Share questions and start discussions in the Mojo 1.0 subcategory of our forum: https://t.co/R2KMzVxYE1 Notice a bug or have a feature request? Open a GitHub Issue with the "Mojo 1.0" label: https://t.co/YWY9xEwodz

"The people who see the most pain are the people writing at the low level and optimizing at the low level. So that's why we love Mojo and MAX - we think that's a way to compete on the same level playing field." - Ramine Roane @roaner, CVP of AI at @AMD, at Full Context, our reception with AMD before their AI DevDay This is the conversation we built these events for. Subscribe to our events calendar: https://t.co/b6mleeuDoE
Available now at https://t.co/rB6zqzBB1g
"A decade ago, AI was supposed to replace radiologists. Today, radiologists make more than $500,000 per year, and their employment continues to grow, see chart below. Reading scans is a task, not a job, and when the task gets cheaper, demand for the job grows." https://t.co/0OszRhslcq
GPT-5.5 Scores .43% on ARC AGI 3! - GPT-5.5: 0.43% - Opus 4.7: 0.18% - GPT-5.4: 0.20% - Claude 4.6: 0.45% - Gemini 3.1: 0.4% The reported failures for GPT 5.5 were: - True local effect, false world model - Wrong level of abstraction from training data - Solved the level, didnβt reinforce the reward I think the full analysis will help OpenAI have a well rounded understanding of where the models are failing in certain modalities
GPT-5.5 & Opus 4.7 on ARC-AGI-3 - GPT-5.5: 0.43% - Opus 4.7: 0.18% We found 3 failure modes: - True local effect, false world model - Wrong level of abstraction from training data - Solved the level, didnβt reinforce the reward See our full analysis π§΅ https://t.co/BYvwvyEFRE
Make sure to read the blog post for a detailed analysis of frontier model failure modes: https://t.co/yZnsjoTA6c
If you want to help the world make sense of AGI and accelerate its arrival, consider joining the ARC Prize foundation. Two roles currently open: Game Platform Engineering Lead, and Model Testing & Analysis Lead https://t.co/xXJmdOWWx0
I wrote Deep Learning with Python to be the definitive guide to how deep learning works and how to best make use of it. Tens of thousands of people got their career start via this book. 120,000 copies sold, and downloaded by millions more. And now it's free to read online: https://t.co/3CbcQ7hmjp
2 Weeks. New Tools. Infinite Worldsπ The World Jam is LIVE. Build the future of interactive 3D with Marble 1.1 + Spark LoD. Join our Discord to start building. More info below π https://t.co/GN3DM60Vil
Really cool weβre at the point where you can build custom, personalized tools for your own niche 3D workflows. Built this with GPT-5.5, Spark, and Marble yesterday to get better precision + control for creating collider meshes for 3DGS experiences. Live: https://t.co/WZnwQYTSC1 https://t.co/TC6KHUra4I
Expand is now available to everyone! Extend your world in any direction you choose: around corners, into rooms, and beyond what you can see π https://t.co/lpJkqgbTNr
60 million Gaussian splats. One massive dark fantasy world ready to explore! βοΈ Created entirely with Marble, this persistent world is brought to life in-browser via our Spark 2.0 LoD system and Three.js Fly through it yourself and learn more about how it was made π https://t.co/VbWfBwiYPS
Summer vibes βοΈππ± Built with Marble, Spark, and Three.js. Persistent World Models let you design for cohesive spaces instead of isolated frames. World Jam ends this weekend! Thereβs still time to build something magical with Marble. More info π https://t.co/wsVFC2S1Tf
We benchmarked Gemma 4 26B-A4B GGUFs to identify the best performing quants. Unsloth ranks first in ALL 22 of 22 model sizes on mean KL divergence, making them SOTA. GGUFs: https://t.co/uOj5RAckXc https://t.co/SUSLMPcUMY

a Japanese dev literally found the ultimate Claude Code cheat code before the rest of us π₯π₯π₯ He simply installs 'Find Skills' and asks: 'Are there any good skills for [GOAL]?' It instantly hands you the perfect match out of 100s π install guide in π§΅β https://t.co/4BDvuYn1p7
Actually useful local models to run on < 1000$ - 24GB vram (4090, 3090) - 24GB Mac Qwen3.6-35B & Gemma-26B = speed Qwen3.5-27B & Gemma-31B = quality Zeta-2 - Cursor tab style Parakeet - Speech to text Hermes-4.3-36B - No Refusals
π Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. π What's new: π§ Outstanding agentic coding β surpasses Qwen3.5-397B-A17B across all major coding benchmarks π‘ Strong reasoning across text & multimodal tasks π Supports thinking & non-thinking modes β Apache 2.0 β fully open, fully yours Smaller model. Bigger results. Community's favorite. β€οΈ We can't wait to see what you build with Qwen3.6-27B! π ππ Blog: https://t.co/P2Zx7FwMxB Qwen Studio: https://t.co/c4vm4LuZrU Github: https://t.co/zKDEbv0R4U Hugging Face: https://t.co/N67hyzxvfr https://t.co/SSdtbWRDap ModelScope: https://t.co/xODf1pj9kw https://t.co/xXhoqlJ2AB

The Local LLM Cheat Sheet for your 32GB RAM device I was asked to put together a practical lineup of local models that fit comfortably on a 32GB machine. At this tier, you start getting access to real flagship-class local models, plus a growing number of custom quants. But for most people, these are the core models worth knowing first. Flagship Models Qwen3.5 27B / GGUF / Q6_K_M The best overall 32GB flagship. General chat, writing, research, and agent workflows. Great if you want one model that can handle almost everything well. Qwen3.6-35B-A3B / GGUF / UD-Q4_K_M Best MoE flagship. Stronger for coding, reasoning, and tool use than most smaller generalists. Gemma 4 31B / GGUF / Q6_K_M Dense premium model. Writing, analysis, reasoning, and high-end local chat. Heavier than the MoE options, but excellent when quality matters more than speed. Models for Fast Flagship Use Gemma 4 26B A4B / GGUF / Q6_K_M Great balance of speed and quality for general assistant work, coding, agent tasks, and research. This is one of the best 32GB picks if you want something that feels high-end without dragging. DeepSeek-R1 Distill Qwen 32B / GGUF / Q4_K_M Offline reasoning engine. Best for math, logic, deliberate analysis, and step-by-step problem solving. Mistral Small 24B / GGUF / Q6_K_M Tool-calling specialist. Strong for assistants, chat workflows, local business tasks, and function calling. Available for 24GB machines. Models for Companion Use Qwen3.5 9B / GGUF / Q6_K_M Best sidekick. Fast drafts, search loops, cheap retries, and secondary agent work. Even on a 32GB machine, you still want a smaller model around for support tasks. Llama 3.1 8B / GGUF / Q6_K_M Long-context companion. RAG, doc ingestion, codebase chat, and long prompts. The output quality is not the sharpest anymore, but it is still useful when needing simple tasks fast. From what my community tells me, the best single models are Qwen3.5 27B or Gemma 4 31B. For two models, the strongest general pairing is Qwen3.5 27B + Qwen3.5 9B. If you are more code-heavy, Qwen3.6-35B-A3B + Llama 3.1 8B. Let me know what models you are running on 32GB, and which ones have actually been worth the RAM.
The Local LLM cheat sheet for your 16GB RAM device I pulled together a lineup of small models that can run comfortably on a Mac Mini or personal laptop while still leaving room for context without melting your machine. Models for Daily Use Qwen3.5 9B / GGUF / Q4_K_M Daily dr
I highly recommend all newcomers to local AI to read my last 2 article before they jump into local AI and purchase any hardware If you're interested in running LLMs and/or diffusion models locally, these will save you a lot of time, pain, and money https://t.co/7mc4RxWYHz
π DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. πΉ DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. πΉ DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today! π Tech Report: https://t.co/drlDrxkYtp π€ Open Weights: https://t.co/T13Y8i7SDM 1/n
