How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
Published in International Conference on Machine Learning (ICML), 2026
Fine-tuning emerging reasoning models on data from a stronger teacher often fails to improve reasoning and can even degrade performance. We identify stylistic divergence between teacher-generated data and the student’s own distribution as a major cause. TESSY interleaves teacher and student models to alternately generate style and non-style tokens, producing sequences that inherit the teacher’s reasoning capabilities while staying stylistically consistent with the student.
