Estimating Underlying Articulatory Targets of Thai Vowels by Using Deep Learning Based on Generating Synthetic Samples from a 3D Vocal Tract Model and Data Augmentation

บทความในวารสาร

ผู้เขียน/บรรณาธิการ

สันติธรรม พรหมอ่อน

กลุ่มสาขาการวิจัยเชิงกลยุทธ์

รายละเอียดสำหรับงานพิมพ์

รายชื่อผู้แต่ง: Lapthawan, Thanat; Prom-On, Santitham; Birkholz, Peter; Xu, Yi;

ผู้เผยแพร่: Institute of Electrical and Electronics Engineers

ปีที่เผยแพร่ (ค.ศ.): 2022

Volume number: 10

หน้าแรก: 41489

หน้าสุดท้าย: 41502

จำนวนหน้า: 14

นอก: 2169-3536

eISSN: 2169-3536

URL: https://www.scopus.com/inward/record.uri?eid=2-s2.0-85128297948&doi=10.1109%2fACCESS.2022.3166922&partnerID=40&md5=bcce6543a7aec8df5cfadbfec05006d5

ภาษา: English-Great Britain (EN-GB)

ดูในเว็บของวิทยาศาสตร์ | ดูบนเว็บไซต์ของสำนักพิมพ์ | บทความในเว็บของวิทยาศาสตร์

บทคัดย่อ

Representation learning is one of the fundamental issues in modeling articulatory-based speech synthesis using target-driven models. This paper proposes a computational strategy for learning underlying articulatory targets from a 3D articulatory speech synthesis model using a bi-directional long short-Term memory recurrent neural network based on a small set of representative seed samples. Using a seeding set from VocalTractLab, a larger training set was generated that provided richer contextual variations for the model to learn. The deep learning model for acoustic-To-Target mapping was then trained to model the inverse relation of the articulation process. This method allows the trained model to map the given acoustic data onto the articulatory target parameters which can then be used to identify the distribution based on linguistic contexts. The model was evaluated based on its effectiveness in mapping acoustics to articulation, and the perceptual accuracy of speech reproduced from the articulation estimated from the recorded speech by native Thai speakers. The model achieved more than 80% phoneme classification accuracy in the listening test conducted with 25 native Thai speakers. The results indicate that the model can accurately imitate speech with a high degree of phonemic precision. © 2013 IEEE.

คำสำคัญ

articulatory model, articulatory target acquisition