Primitive of the day · 2026-09-28
finetune-chat-template-exact-match
A fine-tuned GGUF that loads fine but generates garbage/repetition almost always has a chat-template mismatch — the Ollama Modelfile TEMPLATE must be character-for-character the template the training data used; prevent the whole class by using tokenizer.apply_chat_template() at data-prep time.
When this applies
A fine-tuned model imported into Ollama produces repetition, wrong format, or refusals even though the GGUF loads without error.
Preconditions
- The chat template used to format the training data is known (which token pattern family)
Do this
Set the Modelfile TEMPLATE to exactly the training-time template — per family: Llama-3 <|start_header_id|>/<|eot_id|>, Mistral/Mixtral [INST], Qwen ChatML <|im_start|>/<|im_end|>, Phi-3 <|user|>/<|end|>, Gemma <start_of_turn>/<end_of_turn>; going forward, format training data with apply_chat_template so data, tokenizer, and Modelfile agree.
What you should see
Coherent, correctly formatted generations after ollama create.
How it fails if ignored
The doc's 'MOST COMMON failure mode': garbage output with a perfectly valid GGUF, typically misdiagnosed as a bad fine-tune.
Do not use when
Base models pulled from the Ollama registry (templates ship correct); actual GGUF corruption (re-fuse instead — 'invalid model' at create time).
Kind: gotcha-fix. Part of the skill MLX lm fuse and GGUF export. Free to reuse in your own agent skills.
Get the whole skill
All 7 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skill