In the first phase of the post-training revolution, supervised fine-tuning relied almost exclusively on curated human demonstrations. Research laboratories and enterprise machine learning teams hired armies of software engineers, legal analysts, and domain specialists to write paired question-and-answer datasets, conversational dialogues, and step-by-step reasoning chains. This methodology succeeded in imbuing foundation models with conversational fluency, […]