Speech data augmentation
Abstract
Data augmentation is used for speech emotion recognition tasks where certain emotional labels, e.g., sadness, are significantly underrepresented in a training dataset. This is typical for data collected in real-life applications. We propose conditioned data augmentation using Generative Adversarial Networks (GANs), in order to generate samples for underrepresented emotions. We propose a conditional GAN architecture to generate synthetic spectrograms for the minority class. For comparison purposes, we implement a series of signal-based data augmentation methods. Results on the speech emotion recognition task show that the proposed data augmentation method significantly improves classification performance as compared to traditional speech data augmentation methods.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for configuring a speech processor comprising:
receiving training data include data for a plurality of classes, including first class data and second class data, where the amount of the second class data is substantially less than the amount of the first class data; processing the training data to produce augmented training data, the augmented data including synthesized second class data, wherein the processing of the training data includes configuring a generative adversarial network (GAN) according the the second class data, and using a generator of the GAN to produce the synthesized second class data; and using the augmented training data by a machine-learning training system to determine values of configuration parameters for use by a machine-learning processor.
2 . The method of claim 1 further comprising:
processing an input signals using the machine-learning processor configured with the determined values of the configuration parameters to produce a result, the result characterizing the input signals according to the their similarity to the second class data.
3 . The method of claim 1 wherein the training data comprises spectrogram data representing audio signals acquired using microphones.
4 . The method of claim 1 wherein the classes include emotion-based classes.
5 . The method of claim 1 wherein the classes include constructs representing human internal states and/or traits based classes.
6 . The method of claim 5 wherein the human internal states include at least one of an emotion and a health condition state.
7 . The method of claim 5 wherein the trait based classes include at least one of an identify and a personality class.
8 . The method of claim 5 wherein targets include detection in changes in the construct states.
9 . The method of claim 5 wherein targets include integrated representation of states and traits.
10 . A non-transitory machine-readable medium having instructions stored thereon, wherein the instructions when executed by a computer processor cause a speech processor to be configured, the configuring of the speech processor comprising:
receiving training data include data for a plurality of classes, including first class data and second class data, where the amount of the second class data is substantially less than the amount of the first class data; processing the training data to produce augmented training data, the augmented data including synthesized second class data, wherein the processing of the training data includes configuring a generative adversarial network (GAN) according the the second class data, and using a generator of the GAN to produce the synthesized second class data; and using the augmented training data by a machine-learning training system to determine values of configuration parameters for use by a machine-learning processor.Join the waitlist — get patent alerts
Track US2020335086A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.