Title Synthetic data generation methods and their application to incomplete experimental design
Translation of Title Sintetinių duomenų generavimo metodai ir jų pritaikymas nepilnam eksperimento planui.
Authors Žeruolis, Rokas
Full Text Download
Pages 49
Keywords [eng] defektų aptikimas, generatyviniai priešiškieji tinklai, sintetinių duomenų generavimas, nematuoti sąlyginiai duomenys, laiko eilutės duomenys, Mel spektrograma, fault detection, generative adversarial network, synthetic data generation, unobserved conditional data, time-series data, Mel spectrogram
Abstract [eng] This study aims to address the challenge of incomplete experimental design by developing a framework to generate synthetic data for previously unobserved combinations of discrete operating parameters. In real-world conditions, acquiring vibration signals for every possible combination of gear conditions, torque values, and rotational speeds is often impractical. To overcome this limitation, a framework based on conditional generative adversarial network (cGAN) was employed to generate synthetic Mel spectrograms based on contextual information. The proposed methodology involves preprocessing raw vibration signals by injecting Laplace noise, segmenting them into overlapping windows, finally transforming, signals into Mel spectrograms. These spectrograms are combined with normalized conditional parameters and used to train the GAN models (conditional Wassertein GAN and conditional Wassertein GAN with gradient penalty), which incorporates convolutional neural networks (CNNs) in both the generator and critic architectures. The results demonstrate that the cWGAN-GP outperforms cWGAN, achieving a Pearson Correlation Coefficient (PCC) of 0.840 and Structural Similarity Index (SSIM) of 0.583 after 3000 training epochs, indicating moderate linear similarity and structural resemblance between synthetic and real spectrograms. Moreover, principal component analysis showed good overlap between real and synthetic data. Hovewer, t-SNE vizualisations reveal persistent discrepencies in global data distribution, suggesting selected GAN models limitation of capturing more complex features. The usefulness of the generated synthetic data for fault classification was evaluated using a baseline CNN classifier together with the TSTR technique. Training exclusively on synthetic data yielded poor classification performance (accuracy of 0.178 using cWGAN-GP), while augmenting the real dataset with synthetic samples did not produce meaningful improvements (accuracy of 0.464 using cWGAN) compared to training on real data alone. While the framework successfully generates synthetic data for unobserved conditions, the synthetic data lacks the fidelity required for meaningful improvements in downstream tasks.
Dissertation Institution Vilniaus universitetas.
Type Master thesis
Language English
Publication date 2026