US2017194006A1PendingUtilityA1

Individualized hotword detection models

Assignee: GOOGLE INCPriority: Jul 22, 2015Filed: Mar 17, 2017Published: Jul 6, 2017
Est. expiryJul 22, 2035(~9 yrs left)· nominal 20-yr term from priority
G06F 40/284G10L 15/26G10L 17/04G10L 15/07G10L 17/24G10L 17/18G10L 2015/088G10L 15/22G10L 2015/0638G10L 15/075G06F 16/683G10L 15/16G10L 17/06G10L 15/063G10L 15/02G10L 15/1815G10L 17/08G10L 15/04G10L 13/086
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for presenting notifications in an enterprise system. In one aspect, a method include actions of obtaining enrollment acoustic data representing an enrollment utterance spoken by a user, obtaining a set of candidate acoustic data representing utterances spoken by other users, determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents a similarity between the enrollment acoustic data and the candidate acoustic data, selecting a subset of candidate acoustic data from the set of candidate acoustic data based at least on the similarity scores, generating a detection model based on the subset of candidate acoustic data, and providing the detection model for use in detecting an utterance spoken by the user.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . (canceled) 
     
     
         2 . A computer-implemented method comprising:
 receiving audio data corresponding to a single utterance by a user of a predefined hotword, wherein the predefined hotword is pronounced by the user using a personalized, non-standard pronunciation;   in response to receiving the audio data corresponding to the single utterance by the user of the predefined hotword, downloading audio features corresponding to other users' utterances of the same, predefined hotword in a manner that is indicated as similar to the personalized, non-standard pronunciation;   dynamically generating a hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword; and   using the dynamically generated hotword detection model to detect a likely utterance of the predefined hotword in subsequently received audio data.   
     
     
         3 . The computer implemented method of  claim 2 , comprising:
 during an enrollment process, prompting, by a client device, the user to speak the predefined hotword; and   generating enrollment acoustic data using the received audio data from the user, wherein the audio data comprises the predefined hotword pronounced by the user using the personalized, non-standard pronunciation and additional one or more terms spoken by the user that trigger semantic interpretation of the one or more terms that follow the predefined hotword.   
     
     
         4 . The computer implemented method of  claim 3 , comprising:
 obtaining a set of candidate acoustic data representing utterances that were previously-spoken by the other users, wherein the other users are of a similar type of user to the user.   
     
     
         5 . The computer implemented method of  claim 4 , wherein obtaining the set of candidate acoustic data representing the utterances that were spoken by the other users comprises:
 determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data.   
     
     
         6 . The computer implemented method of  claim 5 , wherein determining the similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data comprises:
 determining a plurality of sub-similarity scores between the enrollment acoustic data and the candidate acoustic data; and   determining the similarity score based on an averaging of the plurality of sub-similarity scores.   
     
     
         7 . The computer implemented method of  claim 2 , wherein dynamically generating the hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword comprises:
 training the hotword detection model to detect the likely utterance of the predefined hotword by the user in the subsequently received audio data corresponding to the single utterance of the predefined hotword by the user and without requiring the user to speak additional utterances of the predefined hotword.   
     
     
         8 . The computer implemented method of  claim 2 , wherein the hotword detection model is based at least on the single utterance and not based on another utterance of the predefined hotword. 
     
     
         9 . A system comprising:
 one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
 receiving audio data corresponding to a single utterance by a user of a predefined hotword, wherein the predefined hotword is pronounced by the user using a personalized, non-standard pronunciation; 
 in response to receiving the audio data corresponding to the single utterance by the user of the predefined hotword, downloading audio features corresponding to other users' utterances of the same, predefined hotword in a manner that is indicated as similar to the personalized, non-standard pronunciation; 
 dynamically generating a hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword; and 
 using the dynamically generated hotword detection model to detect a likely utterance of the hotword in subsequently received audio data. 
   
     
     
         10 . The system of  claim 9 , the operations further comprise:
 during an enrollment process, prompting, by a client device, the user to speak the predefined hotword; and   generating enrollment acoustic data using the received audio data from the user, wherein the audio data comprises the predefined hotword pronounced by the user using the personalized, non-standard pronunciation and additional one or more terms spoken by the user that trigger semantic interpretation of the one or more terms that follow the predefined hotword.   
     
     
         11 . The system of  claim 10 , the operations further comprise:
 obtaining a set of candidate acoustic data representing utterances that were previously-spoken by the other users, wherein the other users are of a similar type of user to the user.   
     
     
         12 . The system of  claim 11 , wherein obtaining the set of candidate acoustic data representing the utterances that were spoken by the other users the operations further comprise:
 determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data.   
     
     
         13 . The system of  claim 12 , wherein determining the similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data the operations further comprise:
 determining a plurality of sub-similarity scores between the enrollment acoustic data and the candidate acoustic data; and   determining the similarity score based on an averaging of the plurality of sub-similarity scores.   
     
     
         14 . The system of  claim 9 , wherein dynamically generating the hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword the operations further comprise:
 training the hotword detection model to detect the likely utterance of the predefined hotword by the user in the subsequently received audio data corresponding to the single utterance of the predefined hotword by the user and without requiring the user to speak additional utterances of the predefined hotword.   
     
     
         15 . The system of  claim 9 , wherein the hotword detection model is based at least on the single utterance and not based on another utterance of the predefined hotword. 
     
     
         16 . A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
 one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
 receiving audio data corresponding to a single utterance by a user of a predefined hotword, wherein the predefined hotword is pronounced by the user using a personalized, non-standard pronunciation; 
 in response to receiving the audio data corresponding to the single utterance by the user of the predefined hotword, downloading audio features corresponding to other users' utterances of the same, predefined hotword in a manner that is indicated as similar to the personalized, non-standard pronunciation; 
 dynamically generating a hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword; and 
 using the dynamically generated hotword detection model to detect a likely utterance of the hotword in subsequently received audio data. 
   
     
     
         17 . The computer-readable medium of  claim 16 , the operations comprising:
 during an enrollment process, prompting, by a client device, the user to speak the predefined hotword; and   generating enrollment acoustic data using the received audio data from the user, wherein the audio data comprises the predefined hotword pronounced by the user using the personalized, non-standard pronunciation and additional one or more terms spoken by the user that trigger semantic interpretation of the one or more terms that follow the predefined hotword.   
     
     
         18 . The computer-readable medium of  claim 17 , the operations comprising:
 obtaining a set of candidate acoustic data representing utterances that were previously-spoken by the other users, wherein the other users are of a similar type of user to the user.   
     
     
         19 . The computer-readable medium of  claim 18 , wherein obtaining the set of candidate acoustic data representing the utterances that were spoken by the other users the operations comprising:
 determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data.   
     
     
         20 . The computer-readable medium of  claim 19 , wherein determining the similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data the operations comprising:
 determining a plurality of sub-similarity scores between the enrollment acoustic data and the candidate acoustic data; and   determining the similarity score based on an averaging of the plurality of sub-similarity scores.   
     
     
         21 . The computer-readable medium of  claim 16 , wherein dynamically generating the hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword the operations comprising:
 training the hotword detection model to detect the likely utterance of the predefined hotword by the user in the subsequently received audio data corresponding to the single utterance of the predefined hotword by the user and without requiring the user to speak additional utterances of the predefined hotword.

Join the waitlist — get patent alerts

Track US2017194006A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.