Six open-weight LLMs that fit a single 24GB GPU at Q4_K_M quantization
Practical local inference floor: 24GB VRAM. A practical guide to Qwen 3.6, Gemma 4, Mistral Small, gpt-oss-20b, and DeepSeek-R1-Distill—each tested for fit, licensing, and use-case fit.
• Qwen 3.6: fast multilingual inference, Apache 2.0
• Gemma 4: instruction-tuned, open weights, Gemma license
• Mistral Small: efficient, permissive license
• gpt-oss-20b: smaller scale, narrow specialization
• DeepSeek-R1-Distill: reasoning-optimized distill from frontier