To improve the naturalness and robustness of synthesized prosody, a autoregressive speech synthesis model based on phoneme-level prosody modeling is proposed. This model enhances prosody modeling from two aspects: inter-word pauses and phoneme durations. To enhance the diversity and accuracy of inter-word pauses, a pause prediction module is proposed at the text frontend. This module predicts multiple pause labels based on the original text, thereby providing accurate references for pause duration modeling in speech synthesis. To enhance the naturalness of phoneme durations, a duration prediction module is proposed. This module predicts a mixture Gaussian distribution for each phoneme and obtains diversified phoneme durations through random sampling. To stabilize phoneme duration modeling in the autoregressive model, an attention-based discrimination module is proposed. This module is applied at each time step of the autoregressive process and avoids alignment disorder through attention and discrimination mechanisms. Experimental results demonstrate that the three proposed modules effectively enhance the naturalness and robustness of prosody modeling, thereby improving the quality of speech synthesis.
CASANOVAE, WEBERJ, SHULBYC, et al. YourTTS: towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone[C]//International Conference on Machine Learning. PMLR, 2022: 2709-2720.
[2]
WANGY X, SKERRY-RYANR, STANTOND,et al .Tacotron:towards end-to-end speech synthesis[EB/OL]. 2017:1703.10135.
[3]
RENY, RUANY J, TANX, et al. FastSpeech:fast,robust and controllable text to speech[EB/OL]. 2019:1905.09263.org/abs/1905.09263v5.
[4]
SHENJ, PANGR M, WEISSR J,et al .Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions[C]//2018 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP). Calgary, AB, Canada.IEEE, 2018: 4779-4783.
[5]
KIMJ, KONGJ, SON J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C]//International Conference on Machine Learning. PMLR, 2021: 5530-5540.
[6]
YANGD, KORIYAMAT, SAITOY, et al .Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech[C]//ICASSP 2023 - 2023 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP). Rhodes Island,Greece. IEEE,2023:1-5.
[7]
DEVLINJ, CHANGM W, LEEK,et al .BERT:pre-training of deep bidirectional transformers for language understanding[EB/OL].2018:1810.04805.org/abs/1810.04805v2.
YUEH J, DUOW X, YANGJ Y. Neighborhood adaptive attention based cross-domain fusion network for speech enhancement[J]. Journal of Hunan University (Natural Sciences), 2023, 50(12): 59-68.(in Chinese)
[10]
BORSOSZ, MARINIERR, VINCENTD,et al. AudioLM:a language modeling approach to audio generation[J]. IEEE/ACM Transactions on Audio,Speech and Language Processing,2023,31:2523-2533.
[11]
MARYN J M S, UMESHS, KATTAS V. S-vectors and TESA:speaker embeddings and a speaker authenticator based on transformer encoder[J]. IEEE/ACM Transactions on Audio,Speech,and Language Processing,2021,30:404-413.
[12]
RADFORDA, WUJ, CHILDR, et al. Language models are unsupervised multitask learners[J]. OpenAI blog, 2019, 1(8): 9.
[13]
BETKERJ .Better speech synthesis through scaling[EB/OL].2023:2305.07243.org/abs/2305.07243v2.
[14]
PAMISETTYG, RAMA MURTY KSRI .Prosody-TTS:an end-to-end speech synthesis system with prosody control[J].Circuits,Systems,and Signal Processing,2023,42(1):361-384.
[15]
MIKOLOVT, CHENK, CORRADOG,et al .Efficient estimation of word representations in vector space[EB/OL]. 2013: 1301.3781.
[16]
CHUNGJ, GULCEHREC, CHOK,et al .Empirical evaluation of gated recurrent neural networks on sequence modeling[EB/OL].2014:1412.3555.org/abs/1412.3555v1.
[17]
MCAULIFFEM, SOCOLOFM, MIHUCS, et al. Montreal forced aligner:trainable text-speech alignment using kaldi[C]//Interspeech 2017. ISCA, 2017: 498-502.
[18]
LANCUCKIA .Fastpitch:parallel text-to-speech with pitch prediction[C]//ICASSP 2021-2021 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP). Toronto, ON, Canada. IEEE, 2021: 6588-6592.
[19]
VAN DEN OORDA, DIELEMANS, ZENH G,et al .WaveNet:a generative model for raw audio[EB/OL].2016:1609.03499.org/abs/1609.03499v2.
[20]
WANGX, TAKAKIS, YAMAGISHIJ,et al .A vector quantized variational autoencoder (VQ-VAE) autoregressive neural $F0$ model for statistical parametric speech synthesis[J].IEEE/ACM Transactions on Audio,Speech,and Language Processing,2019,28:157-170.
[21]
VEAUXC, YAMAGISHIJ, MACDONALDK .SUPERSEDED-CSTR VCTK corpus:English multi-speaker corpus for CSTR voice cloning toolkit[J]. 2016.
[22]
ZENH G, DANGV, CLARKR,et al. LibriTTS: a corpus derived from LibriSpeech for text-to-speech[EB/OL]. 2019: 1904. 02882.
[23]
PRATAPV, XUQ T, SRIRAMA,et al .MLS:a large-scale multilingual dataset for speech research[EB/OL].2020:2012.03411.org/abs/2012.03411v2.
[24]
PANAYOTOVV, CHENG G, POVEYD,et al .Librispeech:an ASR corpus based on public domain audio books[C]//2015 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP). South Brisbane,QLD, Australia. IEEE,2015:5206-5210.
[25]
KONGJ, KIMJ, BAE J. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances in Neural Information Processing Systems, 2020, 33: 17022-17033.
[26]
DUC P, YUK .Phone-level prosody modelling with GMM-based MDN for diverse and controllable speech synthesis[J].IEEE/ACM Transactions on Audio,Speech,and Language Processing,2021, 30: 190-201.
[27]
MIAOC F, LIANGS, CHENM C, et al. Flow-TTS: a non-autoregressive network for text to speech based on flow[C]//ICASSP 2020-2020 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP). Barcelona,Spain. IEEE, 2020: 7209-7213.
[28]
LIB, GULATIA, YUJ H,et al .A better and faster end-to-end model for streaming ASR[C]//ICASSP 2021-2021 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP). Toronto,ON,Canada. IEEE,2021:5634-5638.