Shrinking Giants: Understanding LLM Quantization Models (Q2, Q4, Q6 and Friends)
Why Quantization Matters Large Language Models (LLMs) are huge. Even a “small” 7B parameter model can chew up 14+ GB in FP16 (16-bit floating point). If you’ve tried running one locally without a beefy GPU, you’ve probably noticed your machine crying in pain—or worse, swapping memory like