Your curated collection of saved posts and media
If you want to help the world make sense of AGI and accelerate its arrival, consider joining the ARC Prize foundation. Two roles currently open: Game Platform Engineering Lead, and Model Testing & Analysis Lead https://t.co/xXJmdOWWx0
I wrote Deep Learning with Python to be the definitive guide to how deep learning works and how to best make use of it. Tens of thousands of people got their career start via this book. 120,000 copies sold, and downloaded by millions more. And now it's free to read online: https://t.co/3CbcQ7hmjp
2 Weeks. New Tools. Infinite Worldsπ The World Jam is LIVE. Build the future of interactive 3D with Marble 1.1 + Spark LoD. Join our Discord to start building. More info below π https://t.co/GN3DM60Vil
Really cool weβre at the point where you can build custom, personalized tools for your own niche 3D workflows. Built this with GPT-5.5, Spark, and Marble yesterday to get better precision + control for creating collider meshes for 3DGS experiences. Live: https://t.co/WZnwQYTSC1 https://t.co/TC6KHUra4I
Expand is now available to everyone! Extend your world in any direction you choose: around corners, into rooms, and beyond what you can see π https://t.co/lpJkqgbTNr
60 million Gaussian splats. One massive dark fantasy world ready to explore! βοΈ Created entirely with Marble, this persistent world is brought to life in-browser via our Spark 2.0 LoD system and Three.js Fly through it yourself and learn more about how it was made π https://t.co/VbWfBwiYPS
Summer vibes βοΈππ± Built with Marble, Spark, and Three.js. Persistent World Models let you design for cohesive spaces instead of isolated frames. World Jam ends this weekend! Thereβs still time to build something magical with Marble. More info π https://t.co/wsVFC2S1Tf
We benchmarked Gemma 4 26B-A4B GGUFs to identify the best performing quants. Unsloth ranks first in ALL 22 of 22 model sizes on mean KL divergence, making them SOTA. GGUFs: https://t.co/uOj5RAckXc https://t.co/SUSLMPcUMY

a Japanese dev literally found the ultimate Claude Code cheat code before the rest of us π₯π₯π₯ He simply installs 'Find Skills' and asks: 'Are there any good skills for [GOAL]?' It instantly hands you the perfect match out of 100s π install guide in π§΅β https://t.co/4BDvuYn1p7
Actually useful local models to run on < 1000$ - 24GB vram (4090, 3090) - 24GB Mac Qwen3.6-35B & Gemma-26B = speed Qwen3.5-27B & Gemma-31B = quality Zeta-2 - Cursor tab style Parakeet - Speech to text Hermes-4.3-36B - No Refusals
π Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. π What's new: π§ Outstanding agentic coding β surpasses Qwen3.5-397B-A17B across all major coding benchmarks π‘ Strong reasoning across text & multimodal tasks π Supports thinking & non-thinking modes β Apache 2.0 β fully open, fully yours Smaller model. Bigger results. Community's favorite. β€οΈ We can't wait to see what you build with Qwen3.6-27B! π ππ Blog: https://t.co/P2Zx7FwMxB Qwen Studio: https://t.co/c4vm4LuZrU Github: https://t.co/zKDEbv0R4U Hugging Face: https://t.co/N67hyzxvfr https://t.co/SSdtbWRDap ModelScope: https://t.co/xODf1pj9kw https://t.co/xXhoqlJ2AB

The Local LLM Cheat Sheet for your 32GB RAM device I was asked to put together a practical lineup of local models that fit comfortably on a 32GB machine. At this tier, you start getting access to real flagship-class local models, plus a growing number of custom quants. But for most people, these are the core models worth knowing first. Flagship Models Qwen3.5 27B / GGUF / Q6_K_M The best overall 32GB flagship. General chat, writing, research, and agent workflows. Great if you want one model that can handle almost everything well. Qwen3.6-35B-A3B / GGUF / UD-Q4_K_M Best MoE flagship. Stronger for coding, reasoning, and tool use than most smaller generalists. Gemma 4 31B / GGUF / Q6_K_M Dense premium model. Writing, analysis, reasoning, and high-end local chat. Heavier than the MoE options, but excellent when quality matters more than speed. Models for Fast Flagship Use Gemma 4 26B A4B / GGUF / Q6_K_M Great balance of speed and quality for general assistant work, coding, agent tasks, and research. This is one of the best 32GB picks if you want something that feels high-end without dragging. DeepSeek-R1 Distill Qwen 32B / GGUF / Q4_K_M Offline reasoning engine. Best for math, logic, deliberate analysis, and step-by-step problem solving. Mistral Small 24B / GGUF / Q6_K_M Tool-calling specialist. Strong for assistants, chat workflows, local business tasks, and function calling. Available for 24GB machines. Models for Companion Use Qwen3.5 9B / GGUF / Q6_K_M Best sidekick. Fast drafts, search loops, cheap retries, and secondary agent work. Even on a 32GB machine, you still want a smaller model around for support tasks. Llama 3.1 8B / GGUF / Q6_K_M Long-context companion. RAG, doc ingestion, codebase chat, and long prompts. The output quality is not the sharpest anymore, but it is still useful when needing simple tasks fast. From what my community tells me, the best single models are Qwen3.5 27B or Gemma 4 31B. For two models, the strongest general pairing is Qwen3.5 27B + Qwen3.5 9B. If you are more code-heavy, Qwen3.6-35B-A3B + Llama 3.1 8B. Let me know what models you are running on 32GB, and which ones have actually been worth the RAM.
The Local LLM cheat sheet for your 16GB RAM device I pulled together a lineup of small models that can run comfortably on a Mac Mini or personal laptop while still leaving room for context without melting your machine. Models for Daily Use Qwen3.5 9B / GGUF / Q4_K_M Daily dr
I highly recommend all newcomers to local AI to read my last 2 article before they jump into local AI and purchase any hardware If you're interested in running LLMs and/or diffusion models locally, these will save you a lot of time, pain, and money https://t.co/7mc4RxWYHz
π DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. πΉ DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. πΉ DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today! π Tech Report: https://t.co/drlDrxkYtp π€ Open Weights: https://t.co/T13Y8i7SDM 1/n

Qwen-Image-2.0-Pro is now live ππ Weβve pushed image quality, multilingual text rendering, and instruction following to a new level, while making performance much more consistent across styles.π π Ranked #9 worldwide for Text-to-Image on @arena πTry it now on ModelScope: https://t.co/pPtrbjzzBK https://t.co/raB6WWMEMP APIοΌhttps://t.co/EgYS5qt2bF
Qwen Image 2.0 Pro 2026-04-22 lands at #9 in Text-to-Image Arena. Highlights of the latest image model from @Alibaba_Qwen: - #9 Text-to-Image - #17 Image Edit (Single Image) Top 10 in Text-to-Image categories: - #6 Portraits - #7 Photorealistic & Cinematic Imagery - #7 Art Con

π₯DeepSeek-V4-Pro API is 75% OFF until May 5th, 2026, 15:59 (UTC Time)! Don't miss out on this massive discount. π οΈIntegration Updates: πΉClaude Code: Set model to deepseek-v4-pro[1m] to unlock 1M context! πΉOpenCode: Update to v1.14.24+ πΉOpenClaw: Update to v2026.4.24+ Check the latest official API docs for full details: https://t.co/9J9ZedDpyU
April was a pretty strong month for LLM releases: - Gemma 4 - GLM-5.1 - Qwen3.6 - Kimi K2.6 - DeepSeek V4 All are now added to the LLM Architecture Gallery. More details once I am fully back in May! https://t.co/HDYbWi2pcc
The Local LLM Cheat Sheet for Your 24GB VRAM or Unified-Memory Device Looking for something else? Look for your RAM tier in the QRT thread. At the 24GB level, the trick is that β24GBβ means two very different ceilings. A 4090 with 24GB VRAM is exclusively yours. An M3 Pro or Mac miniβs 24GB is shared with macOS, so each pick has two recommended quants for either option. Best Daily Driving Models Qwen3.6-27B Bleeding-edge dense flagship with a strong balance of math, code, multilingual work, long-context, and hybrid reasoning. If you only want one current-gen 24GB model, this is the headline pick. Mistral-Small-3.2-24B-Instruct-2506 The same Q5 file works on both a 4090 and an M3 Pro 24GB. Fast, practical, and the zero-caveat daily driver for general assistant work, summaries, drafting, and RAG. Gemma-4-31B-it Googleβs dense flagship for writing, multilingual work, instruction following, and summarization. A 4090 gets a near-lossless Q4_1; M3 Pro drops to IQ4_XS and pays a small reasoning tax. Best Reasoning Models Qwen3.6-35B-A3B Newest Qwen MoE flagship with 35B total, ~3B active. Fast, strong for math, logic, step-by-step analysis, and planning. 4090 gets a clean UD-IQ4_NL; M3 Pro is forced to UD-Q3_K_XL, so gpt-oss-20b is the safer unified-memory hard-math pick. gpt-oss-20b OpenAIβs Apache 2.0 MXFP4-native 20B with generous KV-cache headroom. Fast, strong for general reasoning, useful for general chat, long-context, and basic coding, and comfortable on both discrete and unified memory. Gemma-4-26B-A4B-it Sparse MoE with 26B total and 4B active per token. Very fast, good for agents, long sessions, tool use, and general chat. This is the pick when you care about speed and long-running workflows. The bottom line is if you're going with one model then Qwen3.6-27B is there for you. It's incredible at coding as well as all-around reasoning and substance. Which models have you been running on 24GB VRAM or 24GB unified-memory devices? Let me know in the comments.
Local LLM Cheat Sheet Master Collection: All Tiers (April 2026) Bookmark this thread to access the top LLMs for your exact hardware and use case π§΅ https://t.co/mlvWodg7PB
Yeni Siber GΓΌvenlik modelimi finetune'a baΕladΔ±m; Qwen3.6-35B-A3B'yi 1.3 Milyar token'lΔ±k cybersecurity datasetim ile finetune ediyorum π₯ Bu hafta aΓ§Δ±k-kaynak olarak paylaΕacaΔΔ±m. https://t.co/0ifGcR5SSs
Much of the recent content has been to lead up to this Soon https://t.co/4i0uo3UONB
https://t.co/sF6qq5uIXK
Much of the recent content has been to lead up to this Soon https://t.co/4i0uo3UONB
ComfyUI is the most flexible, composable, and powerful open-source media generation tool with a massive ecosystem of workflows and custom nodes. Your Hermes Agent can now install, launch, manage, and run sophisticated @ComfyUI workflows on demand. https://t.co/IXYdAe1Mnf
We implemented @karpathy 's MicroGPT fully on FPGA fabric. No GPU. No PyTorch. No CPU inference loop. Just a transformer burned into hardware, generating 50,000+ tokens/sec. The model is small, but the idea is not: inference does not have to live only in software π https://t.co/FWYGlJOhA6
Weekly best models for your hardware: ~~ 8 to 16gb ~~ Granite models are amazing: [NEW] - https://t.co/uVSUl1scUM Gemma-E4B is a good general QA model - https://t.co/y2xrZKtjWe Qwen3.5-9B is the best at this level imo - https://t.co/WNwxTttMQC ~~ 16 to 64gb ~~ Another larger Granite: This is a general chat model, really dense with world knowledge. [NEW] - https://t.co/aM9Lld3axk - Undisputed kings: The Qwens at various precisions: (Higher ceiling) - https://t.co/oZIWk2iZxA - https://t.co/oZIWk2iZxA The Gemmas at various precisions: (More efficient) - https://t.co/7E9nKmDwGo - https://t.co/fdt9jntwHG ~~ 64 to 128gb ~~ - Ling is a new 100B~ contender decent agent [NEW] https://t.co/FuwQ571a2R - Mistral medium: from my experience their models have been the most consistent! [NEW] https://t.co/GX7Z4zog1y ~~ 128gb - 256gb ~~ Undisputed king: DeepSeek-V4-Flash [NEW] https://t.co/RAf4GQx86c

Weekly best models for your hardware: ~~ 8 to 16gb ~~ Granite models are amazing: [NEW] - https://t.co/uVSUl1scUM Gemma-E4B is a good general QA model - https://t.co/y2xrZKtjWe Qwen3.5-9B is the best at this level imo - https://t.co/WNwxTttMQC ~~ 16 to 64gb ~~ Another larger Granite: This is a general chat model, really dense with world knowledge. [NEW] - https://t.co/aM9Lld3axk - Undisputed kings: The Qwens at various precisions: (Higher ceiling) - https://t.co/oZIWk2iZxA - https://t.co/oZIWk2iZxA The Gemmas at various precisions: (More efficient) - https://t.co/7E9nKmDwGo - https://t.co/fdt9jntwHG ~~ 64 to 128gb ~~ - Ling is a new 100B~ contender decent agent [NEW] https://t.co/FuwQ571a2R - Mistral medium: from my experience their models have been the most consistent! [NEW] https://t.co/GX7Z4zog1y ~~ 128gb - 256gb ~~ Undisputed king: DeepSeek-V4-Flash [NEW] https://t.co/RAf4GQx86c
Introducing open-slide - The slide framework built for agents. Prompt your agent, get a polished deck. $ npx @β open-slide/cli init π https://t.co/bR3WuQtAjl
We made a guide on how to run open LLMs in Claude Code, Codex and OpenClaw. Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp Guide: https://t.co/VienFDSwcg https://t.co/LgyE0hk1E7
We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks. https://t.co/yIG1Ux27Ro
https://t.co/fIMfdfeQbR
The inability of AI models to produce creative variation is a huge gap. The fact that they generate similar ideas limits their ability to do science & the same-y writing limits their usefulness in many other applications This paper showed you can optimize models for creativity https://t.co/37XypFGU8e
Our research, as well as that of other researchers, shows better prompting techniques help a lot, but model training is still a huge limiting factor. https://t.co/Q13pdSH0iG

Oy. According to a new paper in The Lancet, the rate of made-up citations in biomedical papers has increased by more than 12x since 2023. https://t.co/apfYs7l8PJ https://t.co/qAtqYFOTiQ