System and Method for Performing Dual Mode Speech Recognition
Abstract
A system and method is presented for performing dual mode speech recognition, employing a local recognition module on a mobile device and a remote recognition engine on a server device. The system accepts a spoken query from a user, and both the local recognition module and the remote recognition engine perform speech recognition operations on the query, returning a transcription and confidence score, subject to a latency cutoff time. If both sources successfully transcribe the query, then the system accepts the result having the higher confidence score. If only one source succeeds, then that result is accepted. In either case, if the remote recognition engine does succeed in transcribing the query, then a client vocabulary is updated if the remote system result includes information not present in the client vocabulary.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
transmitting a spoken query to a remote recognition system; obtaining a server transcription from the remote recognition system; and responsive to determining that the server transcription contains a word that is missing from a local vocabulary, issuing a request to the server to send a description of the missing word.
2 . The method of claim 1 , further comprising:
receiving the description of the missing word; and updating the local vocabulary with the description of the missing word.
3 . The method of claim 2 , further comprising:
receiving descriptions of words related to the missing word; and updating the local vocabulary with the descriptions of the related words.
4 . The method of claim 3 wherein the missing word and the related words are related to the same topic.
5 . The method of claim 1 , further comprising:
performing local speech recognition on the spoken query, using the local vocabulary; obtaining a server recognition score; and comparing the server recognition score to a local speech recognition score, wherein issuing the request to the server to send a description of the missing word depends on the outcome of the score comparison.
6 . The method of claim 1 , further comprising:
responsive to available memory resources for client vocabulary data being about to run out, performing a garbage collection operation.
7 . The method of claim 1 , further comprising:
combining frequency and recency of use to determine a word frequency amortized over time; and choosing a priority of a non-permanent word using the word frequency amortized over time.
8 . The method of claim 1 , wherein the description includes a phonetic lattice.
9 . A method comprising:
receiving a request from a client device to send a description of a missing word; determining a topic to which the missing word relates; and sending to the client a set of words that relate to the topic and descriptions of the words.
10 . A non-transitory computer readable medium comprising code that, when executed by one or more processors, causes the one or more processors to:
transmit a spoken query to a remote recognition system; obtain a server transcription from the remote recognition system; and responsive to determining that the server transcription contains a word that is missing from a local vocabulary, issue a request to the server to send a description of the missing word.
11 . The non-transitory computer readable medium of claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
receive the description of the missing word; and update the local vocabulary with the description of the missing word.
12 . The non-transitory computer readable medium of claim 11 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
receive descriptions of words related to the missing word; and update the local vocabulary with the descriptions of the related words.
13 . The non-transitory computer readable medium of claim 12 , wherein the missing word and the related words are related to the same topic.
14 . The non-transitory computer readable medium of claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
perform local speech recognition on the spoken query, using the local vocabulary; obtain a server recognition score; and compare the server recognition score to a local speech recognition score, wherein issuing the request to the server to send a description of the missing word depends on the outcome of the score comparison.
15 . The non-transitory computer readable medium of claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
responsive to available memory resources for client vocabulary data being about to run out, perform a garbage collection operation.
16 . The non-transitory computer readable medium of claim 10 , further comprising code that, when executed by one or more processors, causes the one or more processors to:
combine frequency and recency of use to determine a word frequency amortized over time; and choose a priority of a non-permanent word using the word frequency amortized over time.
17 . The non-transitory computer readable medium of claim 10 , wherein the description includes a phonetic lattice.Join the waitlist — get patent alerts
Track US2017256264A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.