A print-ready, 2-page A4 cheat sheet distilled from my local-LLMs video.
Front: one command to a private, free, offline LLM; what a GGUF file is; quantization (16-bit → 4-bit, 4× smaller); the sizing rule worth memorizing — ~1 GB RAM per billion parameters at 4-bit (8 GB → 7B, 16 GB → 13–14B); and the trick most people miss: both tools are OpenAI-compatible servers at localhost:11434.
Back: 6 practical questions — end-to-end setup, what quantization is, sizing models to machines, which model for chat/coding/reasoning (Llama 3.2, Qwen 2.5 Coder, DeepSeek R1 distill), the server trick, and local vs cloud tradeoffs — complete answers, plus 3 traps ("you need a GPU" is the big one).
For anyone building with LLMs on their own hardware. Instant download.