GGUF model loader
Manage local llama.cpp models.
Import GGUF metadata, select the active model, tune generation controls, and prepare server-mode streaming with GPU acceleration and CPU fallback.
Auto detect: CUDA → ROCm → Metal → CPU fallback
mistral-7b-instruct.Q4_K_M.gguf
4.1 GB · Q4_K_M · context 8,192 · Architecture: Mistral · llama.cpp compatible · GPU layers: auto