708 B
708 B
Qwen 3.8 on llama.cpp — deployment notes
Optimal parameters
The optimal parameters for qwen 3.8 on llama.cpp came out of a week of
benchmarks on the homelab GPU. Context length, flash attention and the
batch size matter more than the mmap knob. The full flag set lives in
quadlet/qwen-llamacpp.yaml — the yaml is the source of truth, these
notes are the reasoning behind each picked value.
What changed since last month
Switched the server image to the 0.1.43 release and moved the model files to the NVMe cache drive. Pinned the quant to q4_k_m; the repeat penalty is the knob that kept the rambles honest. The webui now fronts the raw server so chat sessions survive a container restart.