Comment on Knowledge Distillation of Black-Box Large Language Models (2024)Comments−dmezzetti2moWell-Read Students Learn Better: On the Importance of Pre-training Compact ModelsRelated paper that's a good read: https://arxiv.org/abs/1908.08962
Comments
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Related paper that's a good read: https://arxiv.org/abs/1908.08962