Multiple-Stage Knowledge Distillation

Bai, Nanlan; Xu, Chuanyun; Li, Tian; Li, Mengwei; Li, Gang; Zhang, Yang

doi:10.3390/app12199453

Cited by 1 publication

(1 citation statement)

References 37 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…In the realm of multistage knowledge distillation methods, TSKD 23 effectively enhances the testing accuracy of student networks through multistage guidance from teacher networks. OtO 24 employs a joint multistage to multistage training approach between teacher and student networks, achieving significant improvements in multistage knowledge distillation.…”

Section: Related Workmentioning

confidence: 99%

Multistage feature fusion knowledge distillation

Li,

Wang,

et al. 2024

Sci Rep

View full text Add to dashboard Cite

Generally, the recognition performance of lightweight models is often lower than that of large models. Knowledge distillation, by teaching a student model using a teacher model, can further enhance the recognition accuracy of lightweight models. In this paper, we approach knowledge distillation from the perspective of intermediate feature-level knowledge distillation. We combine a cross-stage feature fusion symmetric framework, an attention mechanism to enhance the fused features, and a contrastive loss function for teacher and student models at the same stage to comprehensively implement a multistage feature fusion knowledge distillation method. This approach addresses the problem of significant differences in the intermediate feature distributions between teacher and student models, making it difficult to effectively learn implicit knowledge and thus improving the recognition accuracy of the student model. Compared to existing knowledge distillation methods, our method performs at a superior level. On the CIFAR100 dataset, it boosts the recognition accuracy of ResNet20 from 69.06% to 71.34%, and on the TinyImagenet dataset, it increases the recognition accuracy of ResNet18 from 66.54% to 68.03%, demonstrating the effectiveness and generalizability of our approach. Furthermore, there is room for further optimization of the overall distillation structure and feature extraction methods in this approach, which requires further research and exploration.

show abstract

Section: Related Workmentioning

confidence: 99%