Augmentation of Audiographic Images for Improved Machine Learning
Abstract
Generally, the present disclosure is directed to systems and methods that generate augmented training data for machine-learned models via application of one or more augmentation techniques to audiographic images that visually represent audio signals. In particular, the present disclosure provides a number of novel augmentation operations which can be performed directly upon the audiographic image (e.g., as opposed to the raw audio data) to generate augmented training data that results in improved model performance. As an example, the audiographic images can be or include one or more spectrograms or filter bank sequences.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A computer-implemented method to train a model to perform speech recognition, the method comprising:
obtaining, by one or more computing devices, one or more audiographic images that respectively visually represent one or more audio signals, wherein the audio signals encode one or more human speech utterances; performing, by the one or more computing devices, one or more augmentation operations on each of the one or more audiographic images to generate one or more augmented images; inputting, by the one or more computing devices, the one or more augmented images into a machine-learned audio processing model, wherein the machine-learned audio processing model is configured to perform speech recognition; receiving, by the one or more computing devices, one or more predictions respectively generated by the machine-learned audio processing model based on the one or more augmented images, wherein the one or more predictions comprise textual transcriptions of the one or more human speech utterances; evaluating, by the one or more computing devices, an objective function that scores the one or more predictions respectively generated by the machine-learned audio processing model; and modifying, by the one or more computing devices, respective values of one or more parameters of the machine-learned audio processing model based on the objective function.
22 . The computer-implemented method of claim 21 , wherein the machine-learned audio processing model comprises an encoder model and a decoder model.
23 . The computer-implemented method of claim 21 , wherein the machine-learned audio processing model is configured to generate a series of attention outputs.
24 . The computer-implemented method of claim 21 , wherein the machine-learned audio processing model comprises a sequence to sequence model.
25 . The computer-implemented method of claim 21 , wherein each of the one or more audiographic images is weakly labeled.
26 . The computer-implemented method of claim 21 , wherein performing, by the one or more computing devices, the one or more augmentation operations comprises performing, by the one or more computing devices, a time warping operation on at least one audiographic image of the one or more audiographic images, wherein performing the time warping operation comprises warping image content of the at least one audiographic image along an axis representative of time.
27 . The computer-implemented method of claim 26 , wherein performing the time warping operation comprises fixing spatial dimensions of the at least one audiographic image and warping the image content of the at least one audiographic image to shift a point within the image content a distance along the axis representative of time.
28 . The computer-implemented method of claim 27 , wherein the distance comprises a user-specified hyperparameter or a learned value.
29 . The computer-implemented method of claim 27 , wherein the point within the image content is randomly selected.
30 . The computer-implemented method of claim 21 , wherein performing, by the one or more computing devices, the one or more augmentation operations comprises performing, by the one or more computing devices, a frequency masking operation on at least one audiographic image of the one or more audiographic images, wherein performing the frequency masking operation comprises changing pixel values for image content associated with a certain subset of frequencies represented by the at least one audiographic image.
31 . The computer-implemented method of claim 30 , wherein:
the certain subset of frequencies extends from a first frequency to a second frequency that is spaced a distance from the first frequency; the distance is selected from a distribution extending from zero to a frequency mask parameter.
32 . The computer-implemented method of claim 31 , wherein the frequency mask parameter comprises a user-specified hyperparameter or a learned value.
33 . The computer-implemented method of claim 30 , wherein changing the pixel values for the image content associated with the certain subset of frequencies comprises changing the pixel values for the image content to equal a mean value associated with the at least one audiographic image.
34 . The computer-implemented method of claim 30 , wherein performing the frequency masking operation comprises enforcing an upper bound on a ratio of the certain subset of frequencies to all frequencies.
35 . The computer-implemented method of claim 21 , wherein performing, by the one or more computing devices, the one or more augmentation operations comprises performing, by the one or more computing devices, a time masking operation on at least one audiographic image of the one or more audiographic images, wherein performing the time masking operation comprises changing pixel values for image content associated with a certain subset of a time steps represented by the at least one audiographic image.
36 . The computer-implemented method of claim 35 , wherein:
the certain subset of time steps extends from a first time step to a second time step that is spaced a distance from the first time step; the distance is selected from a distribution extending from zero to a time mask parameter.
37 . The computer-implemented method of claim 21 , wherein the one or more audiographic images comprise:
one or more spectrograms.
38 . A computer system configured to perform operations, the operations comprising:
obtaining, by the computer system, one or more audiographic images that respectively visually represent one or more audio signals, wherein the audio signals encode one or more human speech utterances; performing, by the computer system, one or more augmentation operations on each of the one or more audiographic images to generate one or more augmented images; inputting, by the computer system, the one or more augmented images into a machine-learned audio processing model, wherein the machine-learned audio processing model is configured to perform speech recognition; receiving, by the computer system, one or more predictions respectively generated by the machine-learned audio processing model based on the one or more augmented images, wherein the one or more predictions comprise textual transcriptions of the one or more human speech utterances; evaluating, by the computer system, an objective function that scores the one or more predictions respectively generated by the machine-learned audio processing model; and modifying, by the computer system, respective values of one or more parameters of the machine-learned audio processing model based on the objective function.
39 . The computer system of claim 38 , wherein the machine-learned audio processing model comprises an encoder model and a decoder model, and wherein the machine-learned audio processing model is configured to generate a series of attention outputs, and wherein the machine-learned audio processing model comprises a sequence to sequence model.
40 . One or more non-transitory computer-readable media that store instructions that, when executed by a computer system, cause the computer system to perform operations, the operations comprising:
obtaining, by the computer system, one or more audiographic images that respectively visually represent one or more audio signals, wherein the audio signals encode one or more human speech utterances; performing, by the computer system, one or more augmentation operations on each of the one or more audiographic images to generate one or more augmented images; inputting, by the computer system, the one or more augmented images into a machine-learned audio processing model, wherein the machine-learned audio processing model is configured to perform speech recognition; receiving, by the computer system, one or more predictions respectively generated by the machine-learned audio processing model based on the one or more augmented images, wherein the one or more predictions comprise textual transcriptions of the one or more human speech utterances; evaluating, by the computer system, an objective function that scores the one or more predictions respectively generated by the machine-learned audio processing model; and modifying, by the computer system, respective values of one or more parameters of the machine-learned audio processing model based on the objective function.Join the waitlist — get patent alerts
Track US2023359898A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.