Paper
H2SD: Hybrid Post-hoc Self-Distillation Framework Enhances Large Model Reasoning Capabilities
The Harbin Institute of Technology team proposes a hybrid post-hoc self-distillation framework H2SD, which differentiates the use of teacher signals based on trajectory correctness: for successful trajectories, only the teacher probability is used to adjust the update magnitude, while for failed trajectories, explicit distribution correction is provided through reverse KL divergence. On multiple reasoning benchmarks, H2SD consistently outperforms RLVR, OPSD, and RLSD baselines while maintaining stable optimization and generation efficiency.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT