GGUF model loader

Manage local llama.cpp models.

Import GGUF metadata, select the active model, tune generation controls, and prepare server-mode streaming with GPU acceleration and CPU fallback.

Auto detect: CUDA → ROCm → Metal → CPU fallback

mistral-7b-instruct.Q4_K_M.gguf

4.1 GB · Q4_K_M · context 8,192 · Architecture: Mistral · llama.cpp compatible · GPU layers: auto

Generation controls

Built with GenMB
Built with GenMB