Estimate memory usage for GGUF models based on GPU layers, context length, and cache type. For details, see this blog post.
--gpu-layers in llama.cpp.
--gpu-layers
--ctx-size in llama.cpp.
--ctx-size
Cache quantization.