How I use fine-tuning, quantization and speculative decoding to speed up a voice bot
Why an unquantized 12B won't run on a gaming GPU, how QLoRA and 4-bit quantization make it fit, and how a fine-tuned draft model makes it 1.7x faster, fast enough for a live voice call.