You already saw that fine-tuning flips a text-continuer into a helper. This is how it's done: train the base model on instruction→answer examples, and grade it only on the answer — so it learns to reply, not to ramble.
SFT data isn't raw web text. It's hand-built examples: a prompt, and the exact answer you wish the model would give.
<|user|>, <|assistant|> — tell the model whose turn it is and where the answer begins.Here's what makes SFT work. The model reads the whole sequence and predicts every next token — but the loss, its grade, is counted only on the response.
That masked loss, repeated over hundreds of thousands of pairs, is the whole difference. Same prompt, same knowledge underneath — only the fine-tuning changed.
<|assistant|> marker, a helpful reply should follow.Good SFT needs a lot of clean prompt→answer pairs — and humans are slow and expensive to write them. Next, synthetic data: using strong models to generate and filter training examples at scale.