| Abstract [eng] |
This paper investigates fine-tuning methods for emotional Lithuanian speech synthesis under limited data conditions. An original Lithuanian emotional speech dataset LESS was created, consisting of 960 recordings across three emotional categories: happy, neutral, and sad. Using the VITS architecture, four base models were trained — three using the CVL and LIEPA datasets, and one using the studio-quality ELKG dataset. Based on each base model, twelve emotional speech synthesis models were created using the fine-tuning method. Model quality was evaluated using two methods: objective acoustic analysis comparing synthesized and real speech features, and a subjective MOS study involving 15 native Lithuanian speakers. Results showed that S-group models trained on studio-quality data significantly outperformed B-group models trained on CVL and LIEPA data in both acoustic analysis and subjective evaluation. It was found that acoustic similarity between base model training data and fine-tuning data speakers is an important factor for successful emotional feature transfer. |