@gkisokay
The Local LLM Cheat Sheet for Your 24GB VRAM or Unified-Memory Device Looking for something else? Look for your RAM tier in the QRT thread. At the 24GB level, the trick is that “24GB” means two very different ceilings. A 4090 with 24GB VRAM is exclusively yours. An M3 Pro or Mac mini’s 24GB is shared with macOS, so each pick has two recommended quants for either option. Best Daily Driving Models Qwen3.6-27B Bleeding-edge dense flagship with a strong balance of math, code, multilingual work, long-context, and hybrid reasoning. If you only want one current-gen 24GB model, this is the headline pick. Mistral-Small-3.2-24B-Instruct-2506 The same Q5 file works on both a 4090 and an M3 Pro 24GB. Fast, practical, and the zero-caveat daily driver for general assistant work, summaries, drafting, and RAG. Gemma-4-31B-it Google’s dense flagship for writing, multilingual work, instruction following, and summarization. A 4090 gets a near-lossless Q4_1; M3 Pro drops to IQ4_XS and pays a small reasoning tax. Best Reasoning Models Qwen3.6-35B-A3B Newest Qwen MoE flagship with 35B total, ~3B active. Fast, strong for math, logic, step-by-step analysis, and planning. 4090 gets a clean UD-IQ4_NL; M3 Pro is forced to UD-Q3_K_XL, so gpt-oss-20b is the safer unified-memory hard-math pick. gpt-oss-20b OpenAI’s Apache 2.0 MXFP4-native 20B with generous KV-cache headroom. Fast, strong for general reasoning, useful for general chat, long-context, and basic coding, and comfortable on both discrete and unified memory. Gemma-4-26B-A4B-it Sparse MoE with 26B total and 4B active per token. Very fast, good for agents, long sessions, tool use, and general chat. This is the pick when you care about speed and long-running workflows. The bottom line is if you're going with one model then Qwen3.6-27B is there for you. It's incredible at coding as well as all-around reasoning and substance. Which models have you been running on 24GB VRAM or 24GB unified-memory devices? Let me know in the comments.