US2017256264A1PendingUtilityA1

System and Method for Performing Dual Mode Speech Recognition

Assignee: SOUNDHOUND INCPriority: Nov 18, 2011Filed: May 23, 2017Published: Sep 7, 2017
Est. expiryNov 18, 2031(~5.3 yrs left)· nominal 20-yr term from priority
G10L 2015/0635G10L 17/06G10L 15/30G10L 15/04G10L 15/34G10L 15/08G10L 15/063G10L 2015/081G10L 15/265G10L 15/26
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method is presented for performing dual mode speech recognition, employing a local recognition module on a mobile device and a remote recognition engine on a server device. The system accepts a spoken query from a user, and both the local recognition module and the remote recognition engine perform speech recognition operations on the query, returning a transcription and confidence score, subject to a latency cutoff time. If both sources successfully transcribe the query, then the system accepts the result having the higher confidence score. If only one source succeeds, then that result is accepted. In either case, if the remote recognition engine does succeed in transcribing the query, then a client vocabulary is updated if the remote system result includes information not present in the client vocabulary.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 transmitting a spoken query to a remote recognition system;   obtaining a server transcription from the remote recognition system; and   responsive to determining that the server transcription contains a word that is missing from a local vocabulary, issuing a request to the server to send a description of the missing word.   
     
     
         2 . The method of  claim 1 , further comprising:
 receiving the description of the missing word; and   updating the local vocabulary with the description of the missing word.   
     
     
         3 . The method of  claim 2 , further comprising:
 receiving descriptions of words related to the missing word; and   updating the local vocabulary with the descriptions of the related words.   
     
     
         4 . The method of  claim 3  wherein the missing word and the related words are related to the same topic. 
     
     
         5 . The method of  claim 1 , further comprising:
 performing local speech recognition on the spoken query, using the local vocabulary;   obtaining a server recognition score; and   comparing the server recognition score to a local speech recognition score,   wherein issuing the request to the server to send a description of the missing word depends on the outcome of the score comparison.   
     
     
         6 . The method of  claim 1 , further comprising:
 responsive to available memory resources for client vocabulary data being about to run out, performing a garbage collection operation.   
     
     
         7 . The method of  claim 1 , further comprising:
 combining frequency and recency of use to determine a word frequency amortized over time; and   choosing a priority of a non-permanent word using the word frequency amortized over time.   
     
     
         8 . The method of  claim 1 , wherein the description includes a phonetic lattice. 
     
     
         9 . A method comprising:
 receiving a request from a client device to send a description of a missing word;   determining a topic to which the missing word relates; and   sending to the client a set of words that relate to the topic and descriptions of the words.   
     
     
         10 . A non-transitory computer readable medium comprising code that, when executed by one or more processors, causes the one or more processors to:
 transmit a spoken query to a remote recognition system;   obtain a server transcription from the remote recognition system; and   responsive to determining that the server transcription contains a word that is missing from a local vocabulary, issue a request to the server to send a description of the missing word.   
     
     
         11 . The non-transitory computer readable medium of  claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
 receive the description of the missing word; and   update the local vocabulary with the description of the missing word.   
     
     
         12 . The non-transitory computer readable medium of  claim 11 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
 receive descriptions of words related to the missing word; and   update the local vocabulary with the descriptions of the related words.   
     
     
         13 . The non-transitory computer readable medium of  claim 12 , wherein the missing word and the related words are related to the same topic. 
     
     
         14 . The non-transitory computer readable medium of  claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
 perform local speech recognition on the spoken query, using the local vocabulary;   obtain a server recognition score; and   compare the server recognition score to a local speech recognition score,   wherein issuing the request to the server to send a description of the missing word depends on the outcome of the score comparison.   
     
     
         15 . The non-transitory computer readable medium of  claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
 responsive to available memory resources for client vocabulary data being about to run out, perform a garbage collection operation.   
     
     
         16 . The non-transitory computer readable medium of  claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
 combine frequency and recency of use to determine a word frequency amortized over time; and   choose a priority of a non-permanent word using the word frequency amortized over time.   
     
     
         17 . The non-transitory computer readable medium of  claim 10 , wherein the description includes a phonetic lattice.

Join the waitlist — get patent alerts

Track US2017256264A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.