System and Method for Combining Speech Recognition Outputs From a Plurality of Domain-Specific Speech Recognizers Via Machine Learning
Abstract
Disclosed herein are systems, methods and non-transitory computer-readable media for performing speech recognition across different applications or environments without model customization or prior knowledge of the domain of the received speech. The disclosure includes recognizing received speech with a collection of domain-specific speech recognizers, determining a speech recognition confidence for each of the speech recognition outputs, selecting speech recognition candidates based on a respective speech recognition confidence for each speech recognition output, and combining selected speech recognition candidates to generate text based on the combination.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method comprising:
recognizing, via a processor, a first portion of received speech with a first speech recognizer specific to a first topic domain, to yield a first speech recognition output; recognizing, via the processor, a second portion of the received speech, the second portion being distinct from the first portion, with a second speech recognizer specific to a second topic domain, to yield second speech recognition output; determining confidence scores for the first speech recognition output and the second speech recognition output, to yield a first speech recognition output confidence score and a second speech recognition output confidence score; and generating text by combining, via a machine-learning algorithm, first speech recognition candidates from the first speech recognition output and second speech recognition candidates from the second speech recognition output, wherein the first speech recognition candidates are based on the first speech recognition output confidence score and the second speech recognition candidates are based on the second speech recognition output confidence score.
2 . The method of claim 1 , wherein domains of the first topic and the second topic domain respectively comprise one of travel, banking, and business.
3 . The method of claim 1 , wherein the machine-learning algorithm comprises a mixture of domain-specific speech recognizers from different domains, wherein the mixture of domain-specific speech recognizers comprises two of the following: local business search, web search, Short Messaging Service, question/answering, video search, broadcast news, and voicemail to text.
4 . The method of claim 3 , wherein the combining of the first speech recognition candidates and the second speech recognition candidates further comprises comparing domain-specific speech recognizers in the mixture of domain-specific speech recognizers to select best speech recognition candidates.
5 . The method of claim 1 , wherein combining of the first speech recognition candidates and the second speech recognition candidates further comprises:
dividing the first speech recognition output and the second speech recognition output into substrings; and selecting a best speech recognition candidate for each substring.
6 . The method of claim 1 , further comprising mixing the first speech recognition candidates and the second speech recognition candidates.
7 . The method of claim 1 , further comprising creating a lattice of the first speech recognition candidates and the second speech recognition candidates.
8 . The method of claim 1 , wherein a speech recognition candidate comprises one of a lattice, confidence scores, and speech recognition metadata.
9 . The method of claim 1 , further comprising:
collecting statistics based on the first speech recognition candidates and the second speech recognition candidates; and training parameters associated with the first speech recognizer and the second speech recognizer based on the statistics.
10 . The method of claim 9 , further comprising training the machine-learning algorithm based on the statistics.
11 . The method of claim 10 , wherein the training parameters are based on one of a lattice combination and a neural network graph that learns from an edit distance between the first speech recognition candidates, the second speech recognition candidates, and a correct recognition candidate.
12 . A system comprising:
a processor; and a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:
recognizing, via a processor, a first portion of received speech with a first speech recognizer specific to a first topic domain, to yield a first speech recognition output;
recognizing a second portion of the received speech, the second portion being distinct from the first portion, with a second speech recognizer specific to a second topic domain, to yield second speech recognition output;
determining confidence scores for the first speech recognition output and the second speech recognition output, to yield a first speech recognition output confidence score and a second speech recognition output confidence score; and
generating text by combining, via a machine-learning algorithm, first speech recognition candidates from the first speech recognition output and second speech recognition candidates from the second speech recognition output, wherein the first speech recognition candidates are based on the first speech recognition output confidence score and the second speech recognition candidates are based on the second speech recognition output confidence score.
13 . The system of claim 12 , wherein domains of the first topic and the second topic domain respectively comprise one of travel, banking, and business.
14 . The system of claim 12 , wherein the machine-learning algorithm comprises a mixture of domain-specific speech recognizers from different domains, wherein the mixture of domain-specific speech recognizers comprises two of the following: local business search, web search, Short Messaging Service, question/answering, video search, broadcast news, and voicemail to text.
15 . The system of claim 14 , wherein combining of the first speech recognition candidates and the second speech recognition candidates further comprises comparing domain-specific speech recognizers in the mixture of domain-specific speech recognizers to select best speech recognition candidates.
16 . The system of claim 12 , wherein combining of the speech recognition candidates further comprises:
dividing the first speech recognition output and the second speech recognition output into substrings; and selecting a best speech recognition candidate for each substring.
17 . The system of claim 12 , the computer-readable storage medium having additional instructions stored which result in operations comprising mixing the first speech recognition candidates and the second speech recognition candidates.
18 . The system of claim 12 , the computer-readable storage medium having additional instructions stored which result in operations comprising creating a lattice of the first speech recognition candidates and the second speech recognition candidates.
19 . The system of claim 12 , wherein a speech recognition candidate comprises one of a lattice, confidence scores, and speech recognition metadata.
20 . A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:
recognizing, via a processor, a first portion of received speech with a first speech recognizer specific to a first topic domain, to yield a first speech recognition output; recognizing a second portion of the received speech, the second portion being distinct from the first portion, with a second speech recognizer specific to a second topic domain, to yield second speech recognition output; determining confidence scores for the first speech recognition output and the second speech recognition output, to yield a first speech recognition output confidence score and a second speech recognition output confidence score; and generating text by combining, via a machine-learning algorithm, first speech recognition candidates from the first speech recognition output and second speech recognition candidates from the second speech recognition output, wherein the first speech recognition candidates are based on the first speech recognition output confidence score and the second speech recognition candidates are based on the second speech recognition output confidence score.Join the waitlist — get patent alerts
Track US2014358537A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.