Title Emocionalios lietuvių kalbos sintezė naudojant neuroninius tinklus
Translation of Title TextToSpeech synthesis of emotional lithuanian language based on neural networks.
Authors Marčiukonis, Gintaras
Full Text Download
Pages 52
Abstract [eng] This paper investigates fine-tuning methods for emotional Lithuanian speech synthesis under limited data conditions. An original Lithuanian emotional speech dataset LESS was created, consisting of 960 recordings across three emotional categories: happy, neutral, and sad. Using the VITS architecture, four base models were trained — three using the CVL and LIEPA datasets, and one using the studio-quality ELKG dataset. Based on each base model, twelve emotional speech synthesis models were created using the fine-tuning method. Model quality was evaluated using two methods: objective acoustic analysis comparing synthesized and real speech features, and a subjective MOS study involving 15 native Lithuanian speakers. Results showed that S-group models trained on studio-quality data significantly outperformed B-group models trained on CVL and LIEPA data in both acoustic analysis and subjective evaluation. It was found that acoustic similarity between base model training data and fine-tuning data speakers is an important factor for successful emotional feature transfer.
Dissertation Institution Vilniaus universitetas.
Type Master thesis
Language Lithuanian
Publication date 2026