US2019122666A1PendingUtilityA1
Digital assistant providing whispered speech
Est. expiryJun 10, 2036(~9.9 yrs left)· nominal 20-yr term from priority
G10L 2015/223G10L 13/033G10L 15/22G10L 25/24G10L 17/26G10L 25/18G10L 13/08
48
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and processes for detecting and/or providing a whispered speech response are provided. In one example process, speech is received from a user, and based on the speech input, determined that a whispered speech response is to be provided. Upon determining that a whispered speech response is to be provided, the whispered speech response is generated and provided to the user.
Claims
exact text as granted — not AI-modified1 .- 26 . (canceled)
27 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:
receive a speech input from a user; determine whether providing a whispered speech response is disabled; in accordance with a determination that providing the whispered speech response is not disabled:
determine, based on the speech input, that a whispered speech response is to be provided;
upon determining that a whispered speech response is to be provided, generate the whispered speech response; and
provide the whispered speech response to the user;
in accordance with a determination that providing the whispered speech response is disabled:
generate a non-whispered speech response; and
provide the non-whispered speech response to the user in lieu of the whispered speech response.
28 . The non-transitory computer-readable storage medium of claim 27 , wherein the speech input comprises at least one of an informational request or a request to perform a task.
29 . The non-transitory computer-readable storage medium of claim 28 , wherein the whispered speech response comprises at least one of a response to the informational request or a response associated with performing the task.
30 . The non-transitory computer-readable storage medium of claim 27 , wherein determining whether providing the whispered speech response is disabled is based on a time of day.
31 . The non-transitory computer-readable storage medium of claim 27 , wherein determining whether providing the whispered speech response is disabled is based on a current location of the electronic device.
32 . The non-transitory computer-readable storage medium of claim 27 , wherein determining that the whispered speech response is to be provided comprises at least one of:
determining whether the speech input includes a whispered speech input; and determining whether context data indicates that the whispered speech response is expected.
33 . The non-transitory computer-readable storage medium of claim 32 , wherein determining whether the speech input includes a whispered speech input comprises:
determining whether the speech input includes a whispered speech input using a neural network classifier.
34 . The non-transitory computer-readable storage medium of claim 32 , wherein the whispered speech input is associated with a first spectrum having one or more first spectrum characteristics associated with a whispered speech.
35 . The non-transitory computer-readable storage medium of claim 34 , wherein the one or more first spectrum characteristics comprise at least one of:
a first amplitude, wherein the first amplitude is less than a second amplitude below a threshold frequency, the second amplitude being associated with the non-whispered speech; a first energy, wherein the first energy is less than a second energy below the threshold frequency, the second energy being associated with the non-whispered speech; a first volume, wherein the first volume is less than a second volume by a threshold volume percentage, the second volume being associated with the non-whispered speech; and a first slope of the first spectrum, wherein the first slope of the first spectrum is shifted by a threshold slope percentage with respect to a second slope of the second spectrum, the second slope of the second spectrum being associated with the non-whispered speech.
36 . The non-transitory computer-readable storage medium of claim 34 , wherein determining whether the speech input includes a whispered speech input comprises:
determining whether the speech input includes a whispered speech input using one or more features of the speech input, wherein the one or more features represent one or more spectrum characteristics associated with a spectrum of the speech input.
37 . The non-transitory computer-readable storage medium of claim 36 , wherein determining whether the speech input includes a whispered speech input using the one or more features comprises:
obtaining the spectrum of the speech input; determining the one or more spectrum characteristics associated with the spectrum of the speech input; and determining a first feature and a second feature based on the one or more spectrum characteristics associated with the spectrum of the speech input.
38 . The non-transitory computer-readable storage medium of claim 37 ,
wherein the first feature is a first mel-frequency cepstrum coefficient (MFCC 0 ) representing an energy or an amplitude associated with the spectrum of the speech input; and wherein the second feature is a second mel-frequency cepstrum coefficient (MFCC 1 ) representing a slope associated with the spectrum of the speech input.
39 . The non-transitory computer-readable storage medium of claim 37 , wherein the one or more programs include further instructions for:
obtaining a whisper score based on the first feature to the second feature; and determining whether the whisper score satisfies a score threshold.
40 . The non-transitory computer-readable storage medium of claim 32 , wherein determining whether the context data indicates that the whispered speech response is expected comprises:
obtaining the context data provided by at least one of the electronic device or one or more additional devices communicatively connected to the electronic device; and determining whether the context data satisfy one or more conditions for providing the whispered speech response.
41 . The non-transitory computer-readable storage medium of claim 27 , wherein generating the whispered speech response comprises:
generating an intermediate speech based on the speech input; and generating the whispered speech response using the intermediate speech.
42 . The non-transitory computer-readable storage medium of claim 41 , wherein the intermediate speech has substantially the same content as the whispered speech response.
43 . The non-transitory computer-readable storage medium of claim 41 , wherein generating the intermediate speech based on the speech input comprises:
generating text based on the speech input; performing natural language processing of the text; identifying a user intent based on a result of the natural language processing; and generating the intermediate speech according to the user intent.
44 . The non-transitory computer-readable storage medium of claim 41 , wherein generating the whispered speech response using the intermediate speech comprises:
obtaining a residual signal based on a linear prediction analysis of the intermediate speech; modifying the residual signal; and obtaining the whispered speech response based on a linear prediction synthesis of the modified residual signal.
45 . The non-transitory computer-readable storage medium of claim 44 , wherein obtaining the residual signal based on a linear prediction analysis of the intermediate speech comprises:
obtaining a plurality of speech frames using the intermediate speech; and performing the linear prediction analysis of the plurality of speech frames.
46 . The non-transitory computer-readable storage medium of claim 45 , wherein performing the linear prediction analysis of the plurality of speech frames comprises:
pre-emphasizing the plurality of speech frames; estimating a plurality of linear prediction coefficients; and inverse filtering the pre-emphasized speech frames to obtain the residual signal.
47 . The non-transitory computer-readable storage medium of claim 46 , wherein estimating the plurality of linear prediction coefficients comprises:
performing a windowing on the pre-emphasized plurality of speech frames.
48 . The non-transitory computer-readable storage medium of claim 44 , wherein modifying the residual signal comprises:
receiving a white noise sequence; estimating energy of the white noise sequence and the residual signal; correlating the energy of the white noise sequence and the energy of the residual signal; and compensating the correlated white noise sequence.
49 . The non-transitory computer-readable storage medium of claim 48 , wherein compensating the correlated white noise sequence comprises performing at least one of differentiating, high-pass filtering, or band-pass filtering with respect to the correlated white noise sequence.
50 . The non-transitory computer-readable storage medium of claim 44 , wherein obtaining the whispered speech response based on the linear prediction synthesis of the modified residual signal comprises:
obtaining a plurality of linear prediction coefficients; modifying the linear prediction coefficients; and performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients.
51 . The non-transitory computer-readable storage medium of claim 50 , wherein modifying the linear prediction coefficients comprises:
converting the plurality of linear prediction coefficients to line spectral frequencies; modifying the line spectral frequencies; and generating modified linear prediction coefficients based on the modified line spectral frequencies.
52 . The non-transitory computer-readable storage medium of claim 50 , wherein performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients comprises:
generating a plurality of whispered frames using a synthesis filter and the modified linear prediction coefficients; and generating the whispered speech response using the plurality of whispered frames.
53 . An electronic device, comprising:
one or more processors; memory; and one or more programs stored in the memory, the one or more programs including instructions for:
receiving a speech input from a user;
determining whether providing a whispered speech response is disabled;
in accordance with a determination that providing the whispered speech response is not disabled:
determining, based on the speech input, that a whispered speech response is to be provided;
upon determining that a whispered speech response is to be provided, generating the whispered speech response; and
providing the whispered speech response to the user;
in accordance with a determination that providing the whispered speech response is disabled:
generating a non-whispered speech response; and
providing the non-whispered speech response to the user in lieu of the whispered speech response.
54 . A method for operating a digital assistant, comprising:
at a user device with one or more processors and memory:
receiving a speech input from a user;
determining whether providing a whispered speech response is disabled;
in accordance with a determination that providing the whispered speech response is not disabled:
determining, based on the speech input, that a whispered speech response is to be provided;
upon determining that a whispered speech response is to be provided, generating the whispered speech response; and
providing the whispered speech response to the user;
in accordance with a determination that providing the whispered speech response is disabled:
generating a non-whispered speech response; and
providing the non-whispered speech response to the user in lieu of the whispered speech response.Join the waitlist — get patent alerts
Track US2019122666A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.