System Providing Expressive and Emotive Text-to-Speech
Abstract
A speech to text system includes a text and labels module receiving a text input and providing a text analysis and a label with a phonetic description of the text. A label buffer receives the label from the text and labels module. A parameter generation module accesses the label from the label buffer and generates a speech generation parameter. A parameter buffer receives the parameter from the parameter generation module. An audio generation module receives the text input, the label, and/or the parameter and generates a plurality of audio samples, A scheduler monitors and schedules the text and label module, the parameter generation module, and/or the audio generation module. The parameter generation module is further configured to initialize a voice identifier with a Voice Style Sheet (VSS) parameter, receive an input indicating a modification to the VSS parameter, and modify the VSS parameter according to the modification.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A control interface device configured to produce an output renderable by a speech synthesizer, comprising:
a processor and a memory configured to store non-transient instructions for execution by the processor; a display unit; an input device configured to accept gestures and/or commands to manipulate a graphical object on the display unit, wherein, when executed by the processor, the instructions perform the steps of:
receiving a text string;
associating a voice parameter with a portion of the text string; and
displaying by the display unit the graphical object comprising a representation of the text string and the voice parameter, wherein the voice parameter on the display unit is represented by a visible curve of voice parameter values plotted against frames, and wherein each respective phoneme in the text string is visually associated with particular ones of the frames;
receiving via the input device a command to modify the voice parameter; and
modifying the voice parameter according to the command.
2 . The device of claim 1 , wherein the voice parameter comprises a prosody characteristic.
3 . The device of claim 1 , wherein the voice parameter is bounded by a personality profile consisting of at least one of the group of a vocal tract length, a pitch range, a phrase duration, a pause duration.
4 . The device of claim 1 , wherein modifying the voice parameter of the audio waveform is in accordance with a parameter range of an audio rendering device configured to render the audio waveform.
5 . The device of claim 1 , further comprising the step of converting by the processor a gesture detected by the input device into the command.
6 . The device of claim 1 , further comprising the step of associating a timestamp with the voice parameter.
7 . The device of claim 1 , wherein the display and the input device comprise a touch screen configured to detect a single touch and/or multi-touch gesture.
8 . The device of claim 1 , wherein the voice parameter comprises a markup symbol added to the text string to provide rendering instructions to the speech synthesizer.
9 . The device of claim 8 , wherein the markup symbol indicates a value or range for one or more vocal parameters, selected from the group consisting of pitch, duration, amplitude, vocal tract dimension, sibilance, prosody width, and silence.
10 . The device of claim 9 , wherein the markup symbol indicates the voice parameter is to be randomized to prevent repeated utterances from sounding identical, wherein a degree of randomness is specified by specifying a high and low range for the parameter's value.
11 . The device of claim 1 , wherein the graphical object comprises an envelope controller.
12 . The device of claim 1 wherein the display is configured to present the text string, the voice parameter, and a second voice parameter as a trajectory.
13 . A method for controlling a voice animation for a text-to-speech synthesizer in real-time, comprising the steps of:
initializing a voice identifier with plurality of Voice Style Sheet (VSS) parameters formatted in a VSS file, wherein the VSS file comprises a text string for each respective one of the plurality of VSS parameters that identifies values for an origin, a width, an amplitude, and a sustain for that VSS parameter for rendering text-to-speech; receiving a text string; generating a plurality of phonetic labels for a rendering of the text string; receiving an input indicating a modification to the plurality of VSS parameters; modifying the plurality of VSS parameters according to the modification; and generating audio samples according to the plurality of modified VSS parameters.
14 . The method of claim 13 , wherein the modification refers to a duration of a portion of the voice animation.
15 . The method of claim 13 , wherein the modification refers to an acoustic feature of a portion of the voice animation.
16 . The method of claim 13 , further comprising the step of assigning a timestamp to a voice parameter of the plurality of VSS parameters.
17 . A speech to text system comprising:
a text and labels module configured to receive a text input and provide a text analysis and a label comprising a phonetic description of the text; a label buffer configured to receive the label from the text and labels module; a parameter generation module configured to access the label from the label buffer and generate a speech generation parameter; a parameter buffer configured to receive the parameter from the parameter generation module; an audio generation module configured to receive the text input, the label, and/or the parameter and generate a plurality of audio samples; and a scheduler configured to monitor and schedule at least one of the group consisting of the text and label module, the parameter generation module, and the audio generation module; wherein the parameter generation module is further configured to perform the steps of:
initializing a voice identifier with a Voice Style Sheet (VSS) parameter formatted in a VSS file, wherein the VSS file comprises a text string for each respective one of the plurality of VSS parameters that identifies values for an origin, a width, an amplitude, and a sustain for that VSS parameter for rendering text-to-speech;
receiving an input indicating a modification to the VSS parameter; and
modifying the VSS parameter according to the modification.
18 . The system of claim 17 , further comprising a control interface configured to display the display voice animation control data and provide an interface to receive real-time input to manipulate the animation control data.
19 . The system of claim 17 , wherein the audio generation module further comprises a text-to-speech (TTS) playback device configured to receive input comprising text and formatted control data for rendering by an audio transducer in real-time.
20 . The system of claim 17 , wherein the plurality of audio samples comprises a speech synthesis of the text input.
21 . The device of claim 19 , wherein the audio generation module further comprises an audio transducer.
22 . The device of claim 17 , further comprising a sample buffer configured to receive the plurality of samples from the audio generation module.
23 . The device of claim 1 , wherein the memory stores a plurality of Voice Style Sheet (VSS) parameters formatted in a VSS file, wherein the VSS file comprises a text string for each respective one of the plurality of VSS parameters that identifies values for an origin, a width, an amplitude, and a sustain for that VSS parameter for rendering text-to-speech.
24 . The method of claim 13 , further comprising:
displaying by a display unit, a graphical object comprising a representation of the text string and one of the VSS parameters, wherein the VSS parameter on the display unit is represented by a visible curve of voice parameter values plotted against frames, and wherein each respective phoneme in the text string is visually associated with particular ones of the frames.
25 . The system of claim 17 , further comprising:
a display unit configured to display a graphical object comprising a representation of the text string and one of the plurality of VSS parameters, wherein the VSS parameter on the display unit is represented by a visible curve of the VSS parameter's values plotted against frames, and wherein each respective phoneme in the input text is visually associated with particular ones of the frames.
26 . The device of claim 9 , wherein the markup symbol indicates the voice parameter is to be randomized to prevent repeated utterances from sounding identical, wherein a degree of randomness is specified by specifying a probability that the parameter adjustment will be applied during a current rendering.
27 . A computer-implemented method of statistical parametric speech synthesis, the method comprising:
analyzing a text string with a text analyzer to produce phonetic labels lexically and phonetically describing the text string; using the phonetic labels to access context dependent models for acoustic features and duration; generating parameters from the context dependent models for all of the phonetic labels; translating controls to a voice style sheet format identifying parameters including pitch (fO), spectrum, duration, vocal tract length, and aperiodicity per frame for all the phonetic labels; providing a control interface for real-time manipulation of the parameters; synthesizing a set of audio samples with a vocoder based on the parameters and the real-time manipulation of the parameters at the control interface to produce a synthesized speech waveform of the text string; and rendering the synthesized speech with a rendering system.Join the waitlist — get patent alerts
Track US2022392430A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.