Document-Level Neural TTS Using Curriculum Learning and Attention Masking
Speech synthesis has been developed to the level of natural human-level speech synthesized through an attention-based end-to-end text-to-speech synthesis (TTS) model. However, it is difficult to generate attention when synthesizing a text longer than the trained length or document-level text. In thi...
Saved in:
Main Authors: | Sung-Woong Hwang, Joon-Hyuk Chang |
---|---|
Format: | Article |
Language: | English |
Published: |
IEEE
2021-01-01
|
Series: | IEEE Access |
Subjects: | |
Online Access: | https://ieeexplore.ieee.org/document/9312676/ |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
-
Mask-wearing affects infants' selective attention to familiar and unfamiliar audiovisual speech
by: Lauren N. Slivka, et al.
Published: (2025-02-01) -
Mask Material Filtration Efficiency and Mask Fitting at the Crossroads: Implications during Pandemic Times
by: Karin Ardon-Dryer, et al.
Published: (2021-03-01) -
Investigation of Mask Efficiency for Loose-fitting Masks against Ultrafine Particles and Effect on Airway Deposition Efficiency
by: Wei-Chung Su, et al.
Published: (2021-12-01) -
Approaches and models development of 2013 Curriculum and Merdeka Curriculum
by: Fahira Azzahra, et al.
Published: (2022-12-01) -
The concept of the document in Forensic science
by: V. S. Sezonov
Published: (2022-03-01)