US2022392430A1PendingUtilityA1

System Providing Expressive and Emotive Text-to-Speech

Assignee: D & M HOLDINGS INCPriority: Mar 23, 2017Filed: Aug 3, 2022Published: Dec 8, 2022
Est. expiryMar 23, 2037(~10.7 yrs left)· nominal 20-yr term from priority
G06F 3/017G10L 13/04G06F 3/04883G10L 13/10G10L 13/0335G10L 13/08
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A speech to text system includes a text and labels module receiving a text input and providing a text analysis and a label with a phonetic description of the text. A label buffer receives the label from the text and labels module. A parameter generation module accesses the label from the label buffer and generates a speech generation parameter. A parameter buffer receives the parameter from the parameter generation module. An audio generation module receives the text input, the label, and/or the parameter and generates a plurality of audio samples, A scheduler monitors and schedules the text and label module, the parameter generation module, and/or the audio generation module. The parameter generation module is further configured to initialize a voice identifier with a Voice Style Sheet (VSS) parameter, receive an input indicating a modification to the VSS parameter, and modify the VSS parameter according to the modification.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A control interface device configured to produce an output renderable by a speech synthesizer, comprising:
 a processor and a memory configured to store non-transient instructions for execution by the processor;   a display unit;   an input device configured to accept gestures and/or commands to manipulate a graphical object on the display unit,   wherein, when executed by the processor, the instructions perform the steps of:
 receiving a text string; 
 associating a voice parameter with a portion of the text string; and 
 displaying by the display unit the graphical object comprising a representation of the text string and the voice parameter, wherein the voice parameter on the display unit is represented by a visible curve of voice parameter values plotted against frames, and wherein each respective phoneme in the text string is visually associated with particular ones of the frames; 
 receiving via the input device a command to modify the voice parameter; and 
 modifying the voice parameter according to the command. 
   
     
     
         2 . The device of  claim 1 , wherein the voice parameter comprises a prosody characteristic. 
     
     
         3 . The device of  claim 1 , wherein the voice parameter is bounded by a personality profile consisting of at least one of the group of a vocal tract length, a pitch range, a phrase duration, a pause duration. 
     
     
         4 . The device of  claim 1 , wherein modifying the voice parameter of the audio waveform is in accordance with a parameter range of an audio rendering device configured to render the audio waveform. 
     
     
         5 . The device of  claim 1 , further comprising the step of converting by the processor a gesture detected by the input device into the command. 
     
     
         6 . The device of  claim 1 , further comprising the step of associating a timestamp with the voice parameter. 
     
     
         7 . The device of  claim 1 , wherein the display and the input device comprise a touch screen configured to detect a single touch and/or multi-touch gesture. 
     
     
         8 . The device of  claim 1 , wherein the voice parameter comprises a markup symbol added to the text string to provide rendering instructions to the speech synthesizer. 
     
     
         9 . The device of  claim 8 , wherein the markup symbol indicates a value or range for one or more vocal parameters, selected from the group consisting of pitch, duration, amplitude, vocal tract dimension, sibilance, prosody width, and silence. 
     
     
         10 . The device of  claim 9 , wherein the markup symbol indicates the voice parameter is to be randomized to prevent repeated utterances from sounding identical, wherein a degree of randomness is specified by specifying a high and low range for the parameter's value. 
     
     
         11 . The device of  claim 1 , wherein the graphical object comprises an envelope controller. 
     
     
         12 . The device of  claim 1  wherein the display is configured to present the text string, the voice parameter, and a second voice parameter as a trajectory. 
     
     
         13 . A method for controlling a voice animation for a text-to-speech synthesizer in real-time, comprising the steps of:
 initializing a voice identifier with plurality of Voice Style Sheet (VSS) parameters formatted in a VSS file, wherein the VSS file comprises a text string for each respective one of the plurality of VSS parameters that identifies values for an origin, a width, an amplitude, and a sustain for that VSS parameter for rendering text-to-speech;   receiving a text string;   generating a plurality of phonetic labels for a rendering of the text string;   receiving an input indicating a modification to the plurality of VSS parameters;   modifying the plurality of VSS parameters according to the modification; and   generating audio samples according to the plurality of modified VSS parameters.   
     
     
         14 . The method of  claim 13 , wherein the modification refers to a duration of a portion of the voice animation. 
     
     
         15 . The method of  claim 13 , wherein the modification refers to an acoustic feature of a portion of the voice animation. 
     
     
         16 . The method of  claim 13 , further comprising the step of assigning a timestamp to a voice parameter of the plurality of VSS parameters. 
     
     
         17 . A speech to text system comprising:
 a text and labels module configured to receive a text input and provide a text analysis and a label comprising a phonetic description of the text;   a label buffer configured to receive the label from the text and labels module;   a parameter generation module configured to access the label from the label buffer and generate a speech generation parameter;   a parameter buffer configured to receive the parameter from the parameter generation module;   an audio generation module configured to receive the text input, the label, and/or the parameter and generate a plurality of audio samples; and   a scheduler configured to monitor and schedule at least one of the group consisting of the text and label module, the parameter generation module, and the audio generation module;   wherein the parameter generation module is further configured to perform the steps of:
 initializing a voice identifier with a Voice Style Sheet (VSS) parameter formatted in a VSS file, wherein the VSS file comprises a text string for each respective one of the plurality of VSS parameters that identifies values for an origin, a width, an amplitude, and a sustain for that VSS parameter for rendering text-to-speech; 
 receiving an input indicating a modification to the VSS parameter; and 
 modifying the VSS parameter according to the modification. 
   
     
     
         18 . The system of  claim 17 , further comprising a control interface configured to display the display voice animation control data and provide an interface to receive real-time input to manipulate the animation control data. 
     
     
         19 . The system of  claim 17 , wherein the audio generation module further comprises a text-to-speech (TTS) playback device configured to receive input comprising text and formatted control data for rendering by an audio transducer in real-time. 
     
     
         20 . The system of  claim 17 , wherein the plurality of audio samples comprises a speech synthesis of the text input. 
     
     
         21 . The device of  claim 19 , wherein the audio generation module further comprises an audio transducer. 
     
     
         22 . The device of  claim 17 , further comprising a sample buffer configured to receive the plurality of samples from the audio generation module. 
     
     
         23 . The device of  claim 1 , wherein the memory stores a plurality of Voice Style Sheet (VSS) parameters formatted in a VSS file, wherein the VSS file comprises a text string for each respective one of the plurality of VSS parameters that identifies values for an origin, a width, an amplitude, and a sustain for that VSS parameter for rendering text-to-speech. 
     
     
         24 . The method of  claim 13 , further comprising:
 displaying by a display unit, a graphical object comprising a representation of the text string and one of the VSS parameters, wherein the VSS parameter on the display unit is represented by a visible curve of voice parameter values plotted against frames, and wherein each respective phoneme in the text string is visually associated with particular ones of the frames.   
     
     
         25 . The system of  claim 17 , further comprising:
 a display unit configured to display a graphical object comprising a representation of the text string and one of the plurality of VSS parameters, wherein the VSS parameter on the display unit is represented by a visible curve of the VSS parameter's values plotted against frames, and wherein each respective phoneme in the input text is visually associated with particular ones of the frames.   
     
     
         26 . The device of  claim 9 , wherein the markup symbol indicates the voice parameter is to be randomized to prevent repeated utterances from sounding identical, wherein a degree of randomness is specified by specifying a probability that the parameter adjustment will be applied during a current rendering. 
     
     
         27 . A computer-implemented method of statistical parametric speech synthesis, the method comprising:
 analyzing a text string with a text analyzer to produce phonetic labels lexically and phonetically describing the text string;   using the phonetic labels to access context dependent models for acoustic features and duration;   generating parameters from the context dependent models for all of the phonetic labels;   translating controls to a voice style sheet format identifying parameters including pitch (fO), spectrum, duration, vocal tract length, and aperiodicity per frame for all the phonetic labels;   providing a control interface for real-time manipulation of the parameters;   synthesizing a set of audio samples with a vocoder based on the parameters and the real-time manipulation of the parameters at the control interface to produce a synthesized speech waveform of the text string; and   rendering the synthesized speech with a rendering system.

Join the waitlist — get patent alerts

Track US2022392430A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.