Injecting Text in Self-Supervised Speech Pretraining

Chen, Zhehuai; Zhang, Yu; Rosenberg, Andrew; Ramabhadran, Bhuvana; Wang, Gary; Moreno, Pedro

Computer Science > Computation and Language

arXiv:2108.12226 (cs)

[Submitted on 27 Aug 2021]

Title:Injecting Text in Self-Supervised Speech Pretraining

Authors:Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, Pedro Moreno

View PDF

Abstract:Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretraining from two different modalities: speech and text. The proposed method, tts4pretrain complements the power of contrastive learning in self-supervision with linguistic/lexical representations derived from synthesized speech, effectively learning from untranscribed speech and unspoken text. Lexical learning in the speech encoder is enforced through an additional sequence loss term that is coupled with contrastive loss during pretraining. We demonstrate that this novel pretraining method yields Word Error Rate (WER) reductions of 10% relative on the well-benchmarked, Librispeech task over a state-of-the-art baseline pretrained with wav2vec2.0 only. The proposed method also serves as an effective strategy to compensate for the lack of transcribed speech, effectively matching the performance of 5000 hours of transcribed speech with just 100 hours of transcribed speech on the AMI meeting transcription task. Finally, we demonstrate WER reductions of up to 15% on an in-house Voice Search task over traditional pretraining. Incorporating text into encoder pretraining is complimentary to rescoring with a larger or in-domain language model, resulting in additional 6% relative reduction in WER.

Comments:	submit to ASRU 2021
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
MSC classes:	68T10
ACM classes:	I.2.7
Cite as:	arXiv:2108.12226 [cs.CL]
	(or arXiv:2108.12226v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2108.12226

Submission history

From: Zhehuai Chen [view email]
[v1] Fri, 27 Aug 2021 11:36:40 UTC (188 KB)

Computer Science > Computation and Language

Title:Injecting Text in Self-Supervised Speech Pretraining

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Injecting Text in Self-Supervised Speech Pretraining

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators