AWS Machine Learning Blog 中文 关注 探索基于 Amazon Nova 的监督微调中的自蒸馏推理 在本文中,我们探讨了一种为缺乏推理轨迹的 SFT 定制数据集生成思维 token 的思路。我们首先分析推理抑制问题,随后提出自蒸馏推理(Self-Distilled Reasoning, SDR),并在三个基准测试中进行验证,最后提供实用建议。 Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova aws.amazon.com