Comment on Complement Objective Training with Pytorch LightningparentComments−p1esk5ySo they don’t even try to explain why splitting it into two steps works better? I wonder if anyone tried splitting knowledge distillation optimization like that.
Comments
So they don’t even try to explain why splitting it into two steps works better? I wonder if anyone tried splitting knowledge distillation optimization like that.