arXiv cs.CL论文
使用路由稀疏自编码器解耦语言与副语言信息
针对自监督语音编码器中语言与副语言信息纠缠的问题,提出结合TopK稀疏自编码器、路由特定监督和跨因子对抗训练的方法。在冻结的SPEAR和WavLM编码器上验证,独立探针显示因子特异性保留与抑制:语言信息在语言路由中更强,而说话人身份等副语言因素被有效分离。
这篇是正式发表的长论文,站内提供中文解读,全文请到原文阅读 PDF。
Abstract:Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.