A Phonetic-Semantic Pre-Training Model for Robust Speech Recognition  被引量:1

在线阅读下载全文

作  者:Xueyang Wu Rongzhong Lian Di Jiang Yuanfeng Song Weiwei Zhao Qian Xu Qiang Yang 

机构地区:[1]Department of Computer Science and Engineering,The Hong Kong University of Science and Technology,Hong Kong,China [2]WeBank Co.Ltd.,Shenzhen 518057,China

出  处:《CAAI Artificial Intelligence Research》2022年第1期1-7,共7页CAAI人工智能研究(英文)

摘  要:Robustness is a long-standing challenge for automatic speech recognition(ASR)as the applied environment of any ASR system faces much noisier speech samples than clean training corpora.However,it is impractical to annotate every types of noisy environments.In this work,we propose a novel phonetic-semantic pre-training(PSP)framework that allows a model to effectively improve the performance of ASR against practical noisy environments via seamlessly integrating pre-training,self-supervised learning,and fine-tuning.In particular,there are three fundamental stages in PSP.First,pre-train the phone-to-word transducer(PWT)to map the generated phone sequence to the target text using only unpaired text data;second,continue training the PWT on more complex data generated from an empirical phone-perturbation heuristic,in additional to self-supervised signals by recovering the tainted phones;and third,fine-tune the resultant PWT with real world speech data.We perform experiments on two real-life datasets collected from industrial scenarios and synthetic noisy datasets,which show that the PSP effectively improves the traditional ASR pipeline with relative character error rate(CER)reductions of 28.63%and 26.38%,respectively,in two real-life datasets.It also demonstrates its robustness against synthetic highly noisy speech datasets.

关 键 词:pre-training automatic speech recognition self-supervised learning 

分 类 号:TP391.1[自动化与计算机技术—计算机应用技术]

 

参考文献:

正在载入数据...

 

二级参考文献:

正在载入数据...

 

耦合文献:

正在载入数据...

 

引证文献:

正在载入数据...

 

二级引证文献:

正在载入数据...

 

同被引文献:

正在载入数据...

 

相关期刊文献:

正在载入数据...

相关的主题
相关的作者对象
相关的机构对象