Speech Recognition Accuracy with Natural-Language Understanding based Meta-Speech Systems for Assistant Systems
Abstract
In one embodiment, a method includes receiving, at a client system via a client-side assistant process, a first audio input from a first user. The method includes generating multiple transcriptions corresponding to the first audio input based on multiple client-side automatic speech recognition (ASR) engines. Each ASR engine is associated with a respective domain out of multiple domains. The method includes determining, for each transcription, a combination of one or more tasks and one or more entities to be associated with the transcription. The method includes selecting one or more combinations of tasks and entities from the multiple combinations to be associated with the first audio input based on a determined selection strategy. The method includes presenting, via the client-side assistant process, a response to the first audio input based on the selected combinations.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising, by a client system:
receiving, at the client system via a client-side assistant process, a first audio input from a first user; generating, by the client system, a plurality of transcriptions corresponding to the first audio input based on a plurality of client-side automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a respective domain of a plurality of domains; determining, by the client system, for each transcription, a combination of one or more tasks and one or more entities to be associated with the transcription; selecting, by the client system, one or more combinations of tasks and entities from the plurality of combinations to be associated with the first audio input based on a determined selection strategy; and presenting, at the client system via the client-side assistant process, a response to the first audio input based on the selected combinations.
2 . The method of claim 1 , wherein each ASR engine is associated with one or more agents of a plurality of agents specific to the respective ASR engine.
3 . The method of claim 1 , wherein each domain of the plurality of domains comprises one or more agents specific to the respective domain.
4 . The method of claim 3 , wherein the one or more agents comprise one or more of a first-party agent or a third-party agent.
5 . The method of claim 1 , wherein each domain of the plurality of domains comprises a set of tasks specific to the respective domain.
6 . The method of claim 1 , wherein the plurality of domains are associated with a plurality of agents, and wherein each agent is operable to execute one or more tasks specific to one or more of the domains.
7 . The method of claim 1 , further comprising:
identifying, for each combination of tasks and entities, a domain of the plurality of domains, wherein selecting the one or more combinations of tasks and entities comprises mapping the domain of each combination of tasks and entities to the domain associated with one of the plurality of ASR engines.
8 . The method of claim 7 , wherein the one or more combination of tasks and entities are selected when the domain of the respective combination of tasks and entities matches the domain of one of the plurality of ASR engines.
9 . The method of claim 1 , wherein generating the plurality of transcriptions comprises:
sending the first audio input to each of the ASR engines of the plurality of ASR engines; and receiving the plurality of transcriptions from the plurality of ASR engines.
10 . The method of claim 1 , wherein one or more of the ASR engines of the plurality of ASR engines are third-party ASR engines associated with third-party systems that are separate from and external to the one or more computing systems, the method further comprising:
sending, to one of the third-party ASR engines, the first audio input to generate one or more transcriptions; and receiving, from the one of the third-party ASR engines, the one or more transcriptions generated by the third-party ASR engine, wherein generating the plurality of transcriptions comprises selecting the one or more transcriptions generated from the third-party ASR engine to determine the combination of tasks and entities associated with each respective transcription.
11 . The method of claim 1 , further comprising:
identifying one or more features for each combination of tasks and entities, wherein the one or more features are indicative of whether the combination of tasks and entities have an attribute; and ranking the plurality of combinations based on their respective identified features, wherein selecting the one or more combinations of tasks and entities comprises selecting the one or more combinations of intents and slots based on the ranking of the plurality of combinations.
12 . The method of claim 1 , further comprising:
identifying one or more same combinations of intents and slots from the plurality of combinations; and ranking the one or more same combinations of intents and slots based on the number of same combinations of intents and slots, wherein selecting the one or more combinations of intents and slots comprises using the ranking of the one or more same combinations of intents and slots.
13 . The method of claim 1 , further comprising:
sending the selected combinations to a plurality of agents; receiving a plurality of responses from the plurality of agents corresponding to the selected combinations; ranking the plurality of responses received from the plurality of agents; selecting the response from the plurality of responses based on the ranking of the plurality of responses; and generating the response to the first audio input based on the selected response.
14 . The method of claim 1 , wherein one of the plurality of ASR engines is a combined ASR engine based on two or more discrete ASR engines, and wherein each of the two or more discrete ASR engines is associated with a separate domain of the plurality of domains.
15 . The method of claim 1 , wherein the response comprises one or more of an action to be performed or one or more results generated from a query.
16 . The method of claim 1 , wherein presenting the response comprises presenting a notification of the action to be performed or a list of one or more results.
17 . The method of claim 1 , further comprising:
determining a selection strategy out of a plurality of selection strategies to select the one or more combination of tasks and entities based on a predetermined order of selection strategies.
18 . The method of claim 17 , wherein a first selection strategy of the predetermined order of selection strategies uses an ontology to map the one or more combinations of tasks and entities to a respective ASR engine, and wherein a second selection strategy of the predetermined order of selection strategies is used in response to the first selection strategy failing to map the one or more combinations of tasks and entities to the respective ASR engine.
19 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
receive, at a client system via a client-side assistant process, a first audio input from a first user; generate, by the client system, a plurality of transcriptions corresponding to the first audio input based on a plurality of client-side automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a respective domain of a plurality of domains; determine, by the client system, for each transcription, a combination of one or more tasks and one or more entities to be associated with the transcription; select, by the client system, one or more combinations of tasks and entities from the plurality of combinations to be associated with the first audio input based on a determined selection strategy; and present, at the client system via the client-side assistant process, a response to the first audio input based on the selected combinations.
20 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
receive, at a client system via a client-side assistant process, a first audio input from a first user; generate, by the client system, a plurality of transcriptions corresponding to the first audio input based on a plurality of client-side automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a respective domain of a plurality of domains; determine, by the client system, for each transcription, a combination of one or more tasks and one or more entities to be associated with the transcription; select, by the client system, one or more combinations of tasks and entities from the plurality of combinations to be associated with the first audio input based on a determined selection strategy; and present, at the client system via the client-side assistant process, a response to the first audio input based on the selected combinations.Join the waitlist — get patent alerts
Track US2022327289A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.