AI researcher Nathan Lambert suggests that distillation, particularly in the SFT stage, is becoming less impactful for achieving state-of-the-art model performance, especially as reinforcement learning (RL) techniques scale up. He notes that Chinese models claiming to be 'Claude' likely use small samples for initial SFT, but the bulk of performance comes from extensive post-training recipes involving more SFT data, filtering, and RL. The best teacher models for distillation are often not the cutting-edge ones, complicating the process and potentially benefiting open-weight models.